Every Jev answer comes with a confidence score, and TypeSafe's own advice is sensible: start conservative, test with your own data, adjust. The question nobody answers is how much data.
The tempting answer is "label 50 examples, pick a threshold, ship it". We tested that. With 50 labelled examples, about one in four teams would set a threshold that misses their accuracy target by more than two percentage points - without knowing it.
Here's the test, the numbers, and a method that works.
Why the Threshold Matters#
The standard way to use Jev safely is a confidence gate:
- Above the threshold - Jev's answer is used automatically.
- Below it - the item goes to a fallback: a person, or a bigger model.
Set the threshold too low and wrong answers slip through. Too high and you send work to the fallback that Jev could have handled - wasting the savings you switched for.
The goal is usually phrased as a promise: "Everything Jev handles on its own will be at least 95% accurate." The threshold is how you keep it.
Why You Can't Just Use 0.95#
In our test on 3,080 support messages, Jev's confidence ran high out of the box. Answers with confidence between 0.90 and 0.99 averaged 0.95 confidence but were right only 83% of the time. When we improved the topic descriptions with a few examples, that gap narrowed (0.95 vs 90%), but didn't vanish.
So a threshold of 0.95 doesn't mean "95% accurate". The only way to find the threshold that gives 95% accuracy on your data is to measure it on your data - which means labelling some examples. How many?
The Test#
We used the full results from our Banking77 test: 3,080 real customer support messages across 77 topics, each with Jev's answer, its confidence, and the correct label.
Then we simulated a team setting their threshold, 1,000 times for each sample size:
- Randomly pick n labelled messages (25, 50, 100, 200 or 500).
- Find the lowest threshold where Jev's accuracy on those n messages hits the target (with at least five messages above it).
- Apply that threshold to all the other messages and measure the real accuracy.
Step 3 is what the team would actually experience in production. The difference between what their sample promised and what they got is the risk.
The Results#
Target: 95% accuracy on everything Jev handles on its own, using the better setup (topic names plus three examples each):
| Labelled examples | Real accuracy (typical) | Real accuracy (unlucky 10%) | Missed target by 2+ points | Handled by Jev |
|---|---|---|---|---|
| 25 | 93.2% | 90.1% | 42% of teams | 90% |
| 50 | 94.6% | 90.1% | 26% of teams | 86% |
| 100 | 95.0% | 91.7% | 12% of teams | 84% |
| 200 | 95.4% | 93.0% | 7% of teams | 83% |
| 500 | 95.4% | 93.7% | 4% of teams | 83% |
With the out-of-the-box setup (topic names only), it was worse. The miss rate was 32% at 50 examples and 17% at 200, and Jev could only handle about 58% of messages at the 95% target - because it's less accurate to begin with.
For a stricter 97% target with the better setup: 31% of teams missed with 50 examples, 8% with 200, and under 1% with 500.
What the Numbers Say#
1. Small samples are overconfident too. With 25 or 50 examples, a few lucky correct answers can make a low threshold look safe. The team thinks they're at 95% and ends up at 90%. The typical result looks fine; the unlucky ones are the problem - and you don't know which kind you are.
2. 200 is the practical minimum, 500 is comfortable. At 200 labelled examples, 93 in 100 teams land within two points of their target. At 500, 96 do - and for the 97% target, almost all of them.
3. Better descriptions beat more labels. Adding three examples per topic raised the share Jev could handle at 95% from about 58% to 83%. No amount of threshold tuning gets you that. Fix the question first, then set the threshold.
4. Labelling 200-500 examples is cheaper than it sounds. Someone who knows your categories can label a few hundred short messages in an afternoon or two - especially if they're checking Jev's suggestion rather than starting from scratch. Running them through Jev costs pennies.
A Method That Works#
- Write good criteria first. Name each option clearly and add two or three real examples. Check the most common confusions and fix the descriptions or merge overlapping options.
- Label 200-500 random recent examples. Random matters - hand-picked easy cases will fool you.
- Pick the threshold with a margin. If your target is 95%, choose the threshold that hits 97% on your sample. We tested this too: with 200 labelled examples, aiming for exactly 95% left 43% of teams below 95% in production; aiming for 97% cut that to 8% (and only 1% fell below 93%). The price is that Jev handles a bit less on its own - 73% of messages instead of 83%.
- Use three bands, not two. Auto-act above a high threshold, send the middle band to a quick human check (show Jev's top three suggestions - in our test the right answer was in the top three 96% of the time), and send the lowest band to full manual handling or a bigger model.
- Re-check monthly. Your messages change: new products, new problems, seasonal spikes. Label a fresh random sample of 100 or so each month and confirm you're still on target.
Frequently Asked Questions#
Is Jev's confidence score calibrated?#
Not perfectly out of the box, in our test. Higher confidence reliably meant higher accuracy, but answers around 0.95 confidence were right 83-90% of the time depending on how well the options were described. Treat the score as a ranking, and set the cut-off on your own data.
How many labelled examples do I need to set a Jev threshold?#
In our simulation, 200 examples kept 93% of teams within two points of a 95% target; 500 kept 96%. With 50, roughly a quarter of teams missed by more than two points.
What threshold should I start with?#
There's no universal number - that's the point of this test. In our support-message test, 0.9 gave 96.3% accuracy on 79% of messages with good descriptions, but the right value for your data could be quite different.
Does this apply to other AI models?#
The sampling problem applies to any model with a confidence or probability score. Small labelled samples give noisy estimates, whatever produced the scores.
Data: Banking77 by PolyAI (Casanueva et al., 2020), licensed under CC BY 4.0. Tests run with jev-1.13.0; 1,000 simulated trials per sample size.
Setting thresholds on real data is the core of our Jev Switch Check: we test 500-2,000 of your examples, tune the confidence bands with a safety margin, and hand over the code. Start with a free savings estimate.
AI engineer at BrightBit Digital



