Abstract illustration: one message fanning out to 77 topic dots, most light (correct) and a few dark (wrong), from our Jev routing test
AI Agents

We Ran 3,080 Support Messages Through Jev: Accuracy, Speed, Cost and the Catch

LucaLuca8 min read

Launch numbers are marketing. We wanted our own. So we took a public set of real customer support messages, ran every one through TypeSafe's Jev, and measured what came back: how often it picked the right topic, how fast, what it cost, and where it went wrong.

The short version: out of the box, Jev picked the right one of 77 topics 80.8% of the time. Adding three example messages per topic lifted that to 88.8%. Typical response time was about a quarter of a second, and the whole test cost less than a dollar. The catch: its confidence scores run high until you tune them.

The Test#

The data: Banking77, a public dataset of customer messages sent to a bank's support team, published by PolyAI under a CC BY 4.0 licence. We used the full test set - 3,080 messages - each labelled with one of 77 support topics like "card arrival", "lost or stolen card" or "exchange rate".

77 fine-grained topics is harder than most real support desks, which route to 5-15 queues. Many Banking77 topics also overlap ("order physical card" vs "get physical card"). Think of this as a tough test, not a typical one.

The setup: one Jev choice question per message - "Which support topic does this message belong to?" - with all 77 topics as options. We ran it two ways:

  1. Topic names only - what you'd get if you plugged Jev in with no effort.
  2. Topic names plus three example messages each - the examples came from the dataset's separate training set, so Jev never saw the test messages in advance.

We called jev-latest (which served jev-1.13.0) from the UK, eight requests at a time. Every one of the 6,160 calls succeeded.

The Results#

Topic names onlyNames + 3 examples
Right topic (first answer)80.8%88.8%
Right topic within top 391.2%96.1%
Typical response (median)257 ms273 ms
Slow response (95th percentile)350 ms434 ms
Input tokens per message1,7066,355
Cost per 1,000 messages$0.07$0.27
Failed requests00

Response times include the round trip from the UK to TypeSafe's US servers, so they're what your application would actually see. They sit comfortably inside TypeSafe's claimed 70-500 ms range.

Cost is based on TypeSafe's early-access pricing of $0.042 per million input tokens, with output free. Jev did report about 828 output tokens per call - you just aren't billed for them.

Finding 1: Five Minutes of Setup Buys Eight Points of Accuracy#

The single biggest improvement didn't come from the model. It came from describing the topics better.

With bare names, Jev had to guess what "beneficiary not allowed" or "why verify identity" meant to this particular bank. With three real examples per topic, accuracy jumped from 80.8% to 88.8%, and the right answer landed in the top three 96.1% of the time.

That extra context quadrupled the input size (1,706 → 6,355 tokens), so cost went from $0.07 to $0.27 per 1,000 messages. At 100,000 messages a month, that's the difference between about $7 and $27 a month. Easy trade.

Lesson: whatever model you use for routing, the quality of your category descriptions is the first thing to fix.

Finding 2: Most Mistakes Were Between Near-Identical Topics#

The most common errors with names only:

Actual topicJev pickedTimes
get physical cardchange PIN35
order physical cardget physical card27
why verify identityverify my identity18
beneficiary not allowedfailed transfer18
direct debit not recognisedcard payment not recognised14

Most of these are topics a human agent would also confuse - and in a real support desk, many would go to the same team anyway. With examples added, the top errors spread out and none happened more than 13 times.

If your categories overlap like this, you have two options: merge them, or give Jev examples that show the difference.

Finding 3 (the Catch): Confidence Runs High Out of the Box#

Every Jev answer comes with a confidence score between 0 and 1. The obvious move is "act automatically when confidence is high, send the rest to a human". That works - but the numbers need checking.

ConfidenceShare of messagesActual accuracy (names only)Actual accuracy (with examples)
Below 0.54% / 3%35%41%
0.5 - 0.711% / 7%49%51%
0.7 - 0.914% / 11%63%72%
0.9 - 0.9922% / 17%83%90%
0.99 and above50% / 62%96%98%

With names only, answers scored 0.9-0.99 had an average confidence of 0.95 but were right only 83% of the time. Adding examples narrowed that gap considerably (0.95 confidence, 90% right).

Confidence still does its job - higher really does mean more likely correct - but you can't read 0.95 as "95% sure" without testing on your own data first.

What This Looks Like in a Real Support Desk#

Using the with-examples setup and a simple rule - Jev routes on its own above a confidence threshold, everything else goes to a fallback:

ThresholdRouted automaticallyAccuracy on thoseSent to fallback
None100%88.8%0%
0.790.3%93.2%9.7%
0.979.1%96.3%20.9%
0.9573.4%97.0%26.6%
0.9962.1%98.0%37.9%

At a 0.9 threshold, four in five messages route themselves at 96% accuracy, and one in five goes to a person or a bigger model. For a support desk, that's a very usable split - and because the top-3 accuracy is 96%, even the fallback can be faster if agents see Jev's three suggestions.

How to choose that threshold properly - and how many labelled examples you need to trust it - is a question with a surprisingly firm answer. We'll cover it in a follow-up.

How the Cost Compares#

We didn't run a language model on the same messages in this test, so we won't claim accuracy numbers for one. But the cost side is simple arithmetic.

OpenAI's list prices (September 2026) run from $0.05 per million input tokens (gpt-5-nano) to $10 (gpt-6-astra). Send the same 6,355-token request and input alone costs $0.32 to $63.55 per 1,000 messages, before any output - against $0.27 for Jev.

So the gap depends entirely on which model you'd otherwise use. Against a flagship model it's huge. Against the cheapest models it's small - and because OpenAI discounts repeated prompt text (like our list of 77 topics) through prompt caching, a cheap model can even come out cheaper than Jev on a request like this. Jev's case there rests on its typed answers and confidence scores, not price. We break the numbers down by volume in our Jev pricing guide.

Should You Route Tickets With Jev?#

Yes, if you route high volumes, into categories you can describe clearly, and you can put a fallback behind it for low-confidence answers and outages.

Test first if your categories overlap heavily, your messages contain personal data (Jev is US-hosted - see our UK safety guide), or a wrong route is costly.

Don't bother if your volume is low and speed isn't a problem - we explain why in when not to use Jev.

Frequently Asked Questions#

How accurate is Jev at classifying support tickets?#

In our test on 3,080 banking support messages and 77 topics, Jev was right 80.8% of the time with topic names alone and 88.8% with three example messages per topic. With a 0.9 confidence threshold, it handled 79% of messages on its own at 96.3% accuracy.

How fast is Jev?#

Median response time was 257-273 ms and the 95th percentile was 350-434 ms, measured from the UK including network time.

How much does it cost to classify tickets with Jev?#

$0.07 to $0.27 per 1,000 messages in our test, depending on how much description you include with each topic. Output tokens are currently free.

Can I reproduce this test?#

Yes. Banking77 is public under CC BY 4.0, and the setup is one choice question per message with the 77 topics as options.


Data: Banking77 by PolyAI (Casanueva et al., 2020), licensed under CC BY 4.0. Tests run with jev-1.13.0.

A public dataset tells you what Jev can do. Your own messages tell you what it will do for you. Our Jev Switch Check runs this exact test on your real data, tunes the thresholds, and hands over the code. Start with a free savings estimate.

Luca
Luca

AI engineer at BrightBit Digital

Related articles