Launch numbers are marketing. We wanted our own. So we took a public set of real customer support messages, ran every one through TypeSafe's Jev, and measured what came back: how often it picked the right topic, how fast, what it cost, and where it went wrong.
The short version: out of the box, Jev picked the right one of 77 topics 80.8% of the time. Adding three example messages per topic lifted that to 88.8%. Typical response time was about a quarter of a second, and the whole test cost less than a dollar. The catch: its confidence scores run high until you tune them.
The Test#
The data: Banking77, a public dataset of customer messages sent to a bank's support team, published by PolyAI under a CC BY 4.0 licence. We used the full test set - 3,080 messages - each labelled with one of 77 support topics like "card arrival", "lost or stolen card" or "exchange rate".
77 fine-grained topics is harder than most real support desks, which route to 5-15 queues. Many Banking77 topics also overlap ("order physical card" vs "get physical card"). Think of this as a tough test, not a typical one.
The setup: one Jev choice question per message - "Which support topic does this message belong to?" - with all 77 topics as options. We ran it two ways:
- Topic names only - what you'd get if you plugged Jev in with no effort.
- Topic names plus three example messages each - the examples came from the dataset's separate training set, so Jev never saw the test messages in advance.
We called jev-latest (which served jev-1.13.0) from the UK, eight requests at a time. Every one of the 6,160 calls succeeded.
The Results#
| Topic names only | Names + 3 examples | |
|---|---|---|
| Right topic (first answer) | 80.8% | 88.8% |
| Right topic within top 3 | 91.2% | 96.1% |
| Typical response (median) | 257 ms | 273 ms |
| Slow response (95th percentile) | 350 ms | 434 ms |
| Input tokens per message | 1,706 | 6,355 |
| Cost per 1,000 messages | $0.07 | $0.27 |
| Failed requests | 0 | 0 |
Response times include the round trip from the UK to TypeSafe's US servers, so they're what your application would actually see. They sit comfortably inside TypeSafe's claimed 70-500 ms range.
Cost is based on TypeSafe's early-access pricing of $0.042 per million input tokens, with output free. Jev did report about 828 output tokens per call - you just aren't billed for them.
Finding 1: Five Minutes of Setup Buys Eight Points of Accuracy#
The single biggest improvement didn't come from the model. It came from describing the topics better.
With bare names, Jev had to guess what "beneficiary not allowed" or "why verify identity" meant to this particular bank. With three real examples per topic, accuracy jumped from 80.8% to 88.8%, and the right answer landed in the top three 96.1% of the time.
That extra context quadrupled the input size (1,706 → 6,355 tokens), so cost went from $0.07 to $0.27 per 1,000 messages. At 100,000 messages a month, that's the difference between about $7 and $27 a month. Easy trade.
Lesson: whatever model you use for routing, the quality of your category descriptions is the first thing to fix.
Finding 2: Most Mistakes Were Between Near-Identical Topics#
The most common errors with names only:
| Actual topic | Jev picked | Times |
|---|---|---|
| get physical card | change PIN | 35 |
| order physical card | get physical card | 27 |
| why verify identity | verify my identity | 18 |
| beneficiary not allowed | failed transfer | 18 |
| direct debit not recognised | card payment not recognised | 14 |
Most of these are topics a human agent would also confuse - and in a real support desk, many would go to the same team anyway. With examples added, the top errors spread out and none happened more than 13 times.
If your categories overlap like this, you have two options: merge them, or give Jev examples that show the difference.
Finding 3 (the Catch): Confidence Runs High Out of the Box#
Every Jev answer comes with a confidence score between 0 and 1. The obvious move is "act automatically when confidence is high, send the rest to a human". That works - but the numbers need checking.
| Confidence | Share of messages | Actual accuracy (names only) | Actual accuracy (with examples) |
|---|---|---|---|
| Below 0.5 | 4% / 3% | 35% | 41% |
| 0.5 - 0.7 | 11% / 7% | 49% | 51% |
| 0.7 - 0.9 | 14% / 11% | 63% | 72% |
| 0.9 - 0.99 | 22% / 17% | 83% | 90% |
| 0.99 and above | 50% / 62% | 96% | 98% |
With names only, answers scored 0.9-0.99 had an average confidence of 0.95 but were right only 83% of the time. Adding examples narrowed that gap considerably (0.95 confidence, 90% right).
Confidence still does its job - higher really does mean more likely correct - but you can't read 0.95 as "95% sure" without testing on your own data first.
What This Looks Like in a Real Support Desk#
Using the with-examples setup and a simple rule - Jev routes on its own above a confidence threshold, everything else goes to a fallback:
| Threshold | Routed automatically | Accuracy on those | Sent to fallback |
|---|---|---|---|
| None | 100% | 88.8% | 0% |
| 0.7 | 90.3% | 93.2% | 9.7% |
| 0.9 | 79.1% | 96.3% | 20.9% |
| 0.95 | 73.4% | 97.0% | 26.6% |
| 0.99 | 62.1% | 98.0% | 37.9% |
At a 0.9 threshold, four in five messages route themselves at 96% accuracy, and one in five goes to a person or a bigger model. For a support desk, that's a very usable split - and because the top-3 accuracy is 96%, even the fallback can be faster if agents see Jev's three suggestions.
How to choose that threshold properly - and how many labelled examples you need to trust it - is a question with a surprisingly firm answer. We'll cover it in a follow-up.
How the Cost Compares#
We didn't run a language model on the same messages in this test, so we won't claim accuracy numbers for one. But the cost side is simple arithmetic.
OpenAI's list prices (September 2026) run from $0.05 per million input tokens (gpt-5-nano) to $10 (gpt-6-astra). Send the same 6,355-token request and input alone costs $0.32 to $63.55 per 1,000 messages, before any output - against $0.27 for Jev.
So the gap depends entirely on which model you'd otherwise use. Against a flagship model it's huge. Against the cheapest models it's small - and because OpenAI discounts repeated prompt text (like our list of 77 topics) through prompt caching, a cheap model can even come out cheaper than Jev on a request like this. Jev's case there rests on its typed answers and confidence scores, not price. We break the numbers down by volume in our Jev pricing guide.
Should You Route Tickets With Jev?#
Yes, if you route high volumes, into categories you can describe clearly, and you can put a fallback behind it for low-confidence answers and outages.
Test first if your categories overlap heavily, your messages contain personal data (Jev is US-hosted - see our UK safety guide), or a wrong route is costly.
Don't bother if your volume is low and speed isn't a problem - we explain why in when not to use Jev.
Frequently Asked Questions#
How accurate is Jev at classifying support tickets?#
In our test on 3,080 banking support messages and 77 topics, Jev was right 80.8% of the time with topic names alone and 88.8% with three example messages per topic. With a 0.9 confidence threshold, it handled 79% of messages on its own at 96.3% accuracy.
How fast is Jev?#
Median response time was 257-273 ms and the 95th percentile was 350-434 ms, measured from the UK including network time.
How much does it cost to classify tickets with Jev?#
$0.07 to $0.27 per 1,000 messages in our test, depending on how much description you include with each topic. Output tokens are currently free.
Can I reproduce this test?#
Yes. Banking77 is public under CC BY 4.0, and the setup is one choice question per message with the 77 topics as options.
Data: Banking77 by PolyAI (Casanueva et al., 2020), licensed under CC BY 4.0. Tests run with jev-1.13.0.
A public dataset tells you what Jev can do. Your own messages tell you what it will do for you. Our Jev Switch Check runs this exact test on your real data, tunes the thresholds, and hands over the code. Start with a free savings estimate.
AI engineer at BrightBit Digital



