Skip to main content
ByDiana Kriuchkova
When can your agent use Jev, and when Jev needs to route to a tool like TabPFN? We compare both on detecting fraudulent job postings that mix text, categories and yes/no flags as their giveaways. This task is very familiar in agentic content & marketplace moderation, where fraudulent postings would hurt both users and platform itself. It also serves as a great parallel to more high-stakes tasks, like fraud in payments or anomaly detection. Both models use the max of the available context window & all results use the same 300 test postings. As an outcome, we arrive to the recommendation that you can try testing on your data.

Setup

Run this notebook top to bottom. It downloads the data and runs both APIs; no local data or saved predictions are required. Enter your API keys at the prompts, or set TABPFN_API_KEY and JEV_API_KEY. You can get a TabPFN key at platform.priorlabs.ai.

Download and split

EMSCAD / Fake Job Postings has 17880 postings, including 866 fraud cases (4.8%). The original paper describes that it was annotated by Workday employees. We use all 16 input fields, including descriptions, requirements, industry, employment type and binary flags. Only the identifier and target are excluded. Text, salary (which is provided as string) and missing values pass through unchanged. We hold out 300 job postings as test samples. We exclude any training posting whose features exactly match any test posting. This prevents faulty evaluation from testing samples being provided in the train part.

Read the results

ROC AUC measures how often a fraud case ranks above a legitimate posting: 0.5 is random, 1.0 is perfect. Higher ROC AUC means you’ll get more fraud cases caught - and therefore, matters if you make decisions on the basis of predictions. Brier score is the average squared probability error, and it basically tells us how calibrated models’ predictions are - or in other words, is model over or under confident in it’s predictions. Lower is better.

Plotting helpers

Shared functions for the metrics, ROC curves, comparison bars and calibration plots.

1. Jev

Jev 1.13 has a 32k-token context size for the “state” plus the question. With all 16 fields of the dataset and prompt below, 44 examples is the maximum that fits context window of Jev. These 44 examples contain two fraud cases - same percentage (~4%) of fraudulent samples as the general dataset. As you would with a regular model, Jev receives the same context for each prediction and returns a noul response.
1. Jev Jev reaches ROC AUC 0.8883 from 44 examples. It ranks fraud above legitimate postings better than random. Its Brier score is 0.101; the calibration plots below check whether its probability values match the observed fraud rates.

2. TabPFN-3.5-Plus

The context size limit leaves most examples unused by Jev. TabPFN can use all 17567 eligible training rows here and predict on the same 300 test samples. It could go even higher - to 1M rows - but we just didn’t have that much data here. We use TabPFN-3.5-Plus with eight estimators and seed 42. We do not do anything extra - the data representation is same as for Jev.
2. TabPFN-3.5-Plus 2. TabPFN-3.5-Plus With all available training data, TabPFN reaches ROC AUC 0.9995, compared with 0.8883 for Jev. Brier is 0.007 versus 0.101; lower means smaller probability errors across the test postings.

3. What if you don’t have so much data?

Compare TabPFN with 400 training rows against Jev’s 44-example result. We keep Jev’s 44 examples and add 356 randomly sampled postings from the remaining training pool. All 16 features and the same 300 test postings remain unchanged.
3. What if you don't have so much data? 3. What if you don't have so much data? With 400 training rows, TabPFN reaches ROC AUC 0.9485, versus 0.8883 for Jev with 44 examples. Brier is 0.026 versus 0.101; lower means smaller probability errors.

4. Can you trust Jev’s probabilities?

If postings receive a risk of about 80%, roughly 80% of them should be fraudulent. These calibration plots compare each group’s average predicted risk with its observed fraud rate. Points below the diagonal mean the model overestimates risk; points above it mean it underestimates risk. Labels show the number of postings in each bin; error bars are approximate 95% Wilson intervals for the observed fraud rate. Big intervals are due to low number of predictions within a probability bin.
4. Can you trust Jev's probabilities? As we can see, Jev systematically overestimates risk - mean predicted risk is 27.5%, while 5.0% of postings are fraudulent. Among the 21 postings it assigns 60–80% risk, only 4 are fraud. This basically means that if you were using Jev for decision-making, the agent would ban more innocent posts than fraudulent - and we can predict (pun intended) that your users won’t be happy about that. TabPFN’s mean predicted risks are 4.8% with 400 training rows and 4.4% with the full training set. On the contrast with Jev, TabPFN does not blanket-ban innocent postings - and that means happier users and less support requests to your team (even if the team is agentic!)

Summary

TabPFN is a useful tool for Jev when solving a problem that costs you money, time, user happiness or compute spent. Jev can decide if a problem actually needs TabPFN - in the fraud example we’ve used here, less fraud means better user experience - and happier customers. In other examples, like fraud in financial industry, fraud might cost you a lot of money and legal headache. In this cookbook, TabPFN has shown a good advantage at a problem where a lot of signal actually comes from text. On the larger context size, the ROC AUC (how well the model ranks fraud) was higher compared with Jev’s 0.8883. With 400 training rows, ROC AUC was higher: 0.9485 versus 0.8883. It also had a way better Brier score. The calibration plots show how much Jev overestimates fraud risk on this test set. If you don’t want to trust our word for it, try Jev + TabPFN by getting an API key at https://platform.priorlabs.ai and giving both to your agent - and let us know how it goes!