Abstract chart: one tall orange bar for Jev towering over four shorter bars for GPT, Claude and Gemini, showing decisions per second
AI Agents

Jev vs GPT and Claude: What the Independent Benchmarks Actually Show

DomDom8 min read

"Is Jev better than GPT?" is the question everyone asks, and most answers so far repeat TypeSafe's launch claims: up to 193x faster and 445x cheaper. TypeSafe built those numbers on four of its own workflows and says they're probably at the high end.

So what happens in an independent test, on public data, with the answers checked against real labels?

We've tested Jev ourselves on 8,654 messages - see our support routing test and spam test. The head-to-head numbers below come from an independent benchmark published by another automation agency on 19 September 2026. We've summarised their results and checked them against our own Jev data - and the two line up closely.

The short version: Jev was about as accurate as small GPT, Claude and Gemini models, two to three and a half times faster, and several times cheaper. A bigger model was more accurate at routing - but Jev with a confidence gate and that bigger model as a fallback matched the bigger model's accuracy at under a third of its cost.

How the Independent Test Worked#

  • Five systems: Jev 1.13, GPT-5.4 nano, Gemini 3.5 Flash-Lite, Claude Haiku 4.5 and GPT-5.6 Terra.
  • Three tasks, all on public data:
    • 8-way routing: 160 bank customer messages, 8 card-related topics (Banking77).
    • 77-way routing: 231 messages across all 77 Banking77 topics.
    • Prompt-injection detection: 400 messages, 170 of them attempts to hijack an AI's instructions.
  • Same conditions for every model: temperature 0, answers forced into the list of allowed labels, reasoning set to minimal, and topic names only - no descriptions or examples.
  • Scale: 791 labelled decisions, 3,955 calls in total, one run per system.

Accuracy#

Each result has a 95% confidence interval in brackets - the range the true accuracy likely falls in, given the sample size.

System8-way routing77-way routingInjection detection
Jev 1.1383.8% (77.3-88.7)78.8% (73.1-83.6)87.0% (83.3-89.9)
GPT-5.4 nano90.0% (84.4-93.8)78.4% (72.6-83.2)80.8% (76.6-84.3)
Gemini 3.5 Flash-Lite77.5% (70.4-83.3)78.4% (72.6-83.2)86.8% (83.1-89.7)
Claude Haiku 4.578.8% (71.8-84.4)76.2% (70.3-81.2)89.0% (85.6-91.7)
GPT-5.6 Terra89.4% (83.6-93.3)84.0% (78.7-88.2)86.8% (83.1-89.7)

What it says: no system won everything. Jev sat in the pack with the small models - ahead on some tasks, behind on others - and most gaps are smaller than the confidence intervals, so they could be noise. The clearest result is GPT-5.6 Terra, the largest model tested, leading Jev by about five points on 77-way routing.

Speed#

SystemMedian95th percentile
Jev 1.130.33 s0.44 s
Gemini 3.5 Flash-Lite0.67 s0.89 s
Claude Haiku 4.51.02 s1.95 s
GPT-5.4 nano1.15 s2.07 s
GPT-5.6 Terra1.17 s2.54 s

Jev was 2.0 to 3.6 times faster at the median - and more consistent, with its slow responses still under half a second. That gap matters most in real-time work, like a voice agent routing a call while the caller waits.

Cost per 1,000 Decisions#

System8-way77-wayInjection
Jev 1.13$0.015$0.040$0.014
GPT-5.4 nano$0.070$0.205$0.081
Gemini 3.5 Flash-Lite$0.087$0.301$0.086
Claude Haiku 4.5$0.357$1.255$0.344
GPT-5.6 Terra$0.609$1.957$0.673

Against the two cheapest models tested, Jev was 4.7 to 7.5 times cheaper. Against Claude Haiku and GPT-5.6 Terra, the gap was far larger.

One caveat from our side: cheaper models exist than the ones tested. OpenAI's gpt-5-nano, for example, lists at $0.05 per million input tokens - close to Jev's $0.042. Against the very cheapest options, the cost gap shrinks a lot; we break that down in our Jev pricing guide.

The Most Useful Result: Jev Plus a Fallback#

The benchmark also tested the setup we recommend to clients: let Jev answer when it's confident, and send everything else to a bigger model.

With a confidence gate of 0.80 and GPT-5.6 Terra as the fallback:

TaskSent to the bigger modelAccuracy of the combinationTerra aloneCost vs Terra alone
8-way routing19%90.0%89.4%26%
77-way routing23%84.8%~84%28%

The combination was at least as accurate as the big model on its own, for roughly a quarter of the cost - and with an average response of about 0.7-0.8 seconds, still faster.

This is the real answer to "Jev vs GPT": for decisions, it's often not either/or. Jev handles the confident majority; a language model handles the rest.

How It Compares With Our Own Tests#

We ran Jev on far more messages - the full 3,080-message Banking77 test set - so we can check whether the independent Jev numbers hold up. They do:

Jev, 77-way routing, topic names onlyIndependent test (231 messages)Our test (3,080 messages)
Accuracy78.8%80.8%
At a 0.9 confidence gate: share answered66.2%71.0%
At a 0.9 confidence gate: accuracy92.2%92.1%
Median response time0.33 s0.26 s

Two independent runs, different sample sizes, very similar results. That gives us more confidence in both.

Our test adds one thing theirs didn't try: describing the topics better. Adding three example messages per topic lifted Jev from 80.8% to 88.8% on the 77-way task - above every model's 77-way result in the independent test (best: 84.0%). But the language models weren't given examples either, and they would likely improve too. That's a comparison nobody has published yet.

The Limits of What We Know#

This is the best head-to-head data available so far, but it's early:

  • Small samples. The confidence intervals are 6-13 points wide. Differences smaller than that need more data.
  • One run, public data. Banking77 and the injection dataset may have appeared in the language models' training data.
  • Labels only. No model was given descriptions of the options. Real deployments usually include them, which could change the ranking.
  • Latency depends on setup. Their times came from one laptop with four parallel requests; ours from the UK with eight. Your numbers will differ.
  • No frontier models. The biggest model tested was GPT-5.6 Terra, not flagships like GPT-6 Astra or Claude Opus.

So, Should You Use Jev or GPT?#

Your situationOur take
High-volume decisions, speed mattersJev, with a confidence gate
Accuracy is critical and you're on a big modelJev first, big model as fallback - similar accuracy, a fraction of the cost
You need text, reasoning or explanationsA language model - Jev doesn't write text
Low volume on a cheap modelProbably not worth switching - see when not to use Jev
Personal data, UK businessCheck the rules first - see our UK safety guide

Frequently Asked Questions#

Is Jev more accurate than GPT?#

Not in general. In the independent benchmark, Jev was roughly as accurate as small models like GPT-5.4 nano and Claude Haiku 4.5, and about five points behind the larger GPT-5.6 Terra on 77-way routing. Most gaps were within the margin of error.

Is Jev faster than GPT and Claude?#

Yes, in the independent test - 2.0 to 3.6 times faster at the median, at 0.33 seconds versus 0.67 to 1.17 seconds for the language models tested. In our own test, Jev's median was about a quarter of a second.

Is Jev cheaper than Claude?#

Much cheaper than Claude Haiku 4.5 in the benchmark: $0.04 versus $1.26 per 1,000 decisions on 77-way routing. Against the very cheapest language models the gap is much smaller.

Can Jev replace GPT or Claude?#

For typed decisions like routing, scoring and yes/no checks, it can replace many calls. For writing, reasoning or explaining, no. The strongest setup in the benchmark used both: Jev for confident answers, a language model for the rest.#

Public benchmarks tell you what's possible. Your own data tells you what you'll get. Our Jev Switch Check runs this comparison on your real AI calls - Jev, a backup model and your current setup - and shows you the accuracy, speed and cost of each. Start with a free savings estimate.

Dom
Dom

AI engineer at BrightBit Digital

Related articles