If your website has a contact form, you know the problem: real enquiries buried under SEO pitches, fake offers and bot junk. Someone has to sort it, and every minute they spend on spam is a minute a genuine lead waits.
A spam check is one of the simplest jobs you can give TypeSafe's Jev - a single yes/no question per message. So we tested it on a public dataset of 5,574 real messages, each already labelled spam or genuine, and measured exactly what it caught and what it missed.
The short version: Jev separated spam from genuine messages very well (99.2% on the standard ranking measure). At a strict setting it caught 84% of spam while wrongly flagging just 9 of 4,827 genuine messages. It cost $0.015 per 1,000 messages and answered in about a quarter of a second.
The Test#
The data: the SMS Spam Collection from the UCI Machine Learning Repository, licensed CC BY 4.0. It contains 5,574 real text messages: 747 spam and 4,827 genuine.
We'll be upfront: these are text messages, not website enquiries. There's no public, labelled dataset of contact-form submissions - businesses don't publish their inboxes. But the job is the same: spot unsolicited promotions, scams and bait among genuine messages. Treat these results as a strong indication, then confirm on your own form data.
The setup: one Jev noul question per message - the probability that a statement is true:
Is this message spam: unsolicited marketing, a scam, or prize or offer bait, rather than a genuine message from someone the recipient knows or deals with?
Jev returns a number from 0 to 1. We ran all 5,574 messages through jev-latest (serving jev-1.13.0) from the UK. Every call succeeded.
The Results#
Ranking quality (AUC): 0.992. In plain English: pick one spam and one genuine message at random, and Jev gives the spam a higher score 99.2% of the time. That's the headline measure of how well a score separates two groups, independent of where you draw the line.
Where you draw the line decides the trade-off:
| Block if score is at least | Spam caught | Of blocked messages, actually spam | Genuine messages wrongly blocked |
|---|---|---|---|
| 0.5 | 96.3% | 89.5% | 84 (1.74%) |
| 0.7 | 92.8% | 94.9% | 37 (0.77%) |
| 0.8 | 90.1% | 96.6% | 24 (0.50%) |
| 0.9 | 84.3% | 98.6% | 9 (0.19%) |
| 0.95 | 75.5% | 99.8% | 1 (0.02%) |
Speed: median 248 ms, 95th percentile 355 ms, including the round trip to TypeSafe's US servers.
Cost: about 360 input tokens per message - $0.015 per 1,000 messages at TypeSafe's early-access price, with output free. The entire 5,574-message test cost under ten cents.
Don't Pick One Line - Use Three Bands#
For a contact form, a missed genuine enquiry costs far more than a spam message that gets through. So a single "block" threshold is the wrong design. Use three bands instead:
| Band | What happens | Messages | Mistakes |
|---|---|---|---|
| Score 0.9 and above | Straight to the spam folder | 639 | 9 genuine messages among them |
| Score 0.3 - 0.9 | Quick human check | 332 (6%) | - |
| Below 0.3 | Straight to the inbox | 4,603 | 16 spam got through |
With this setup, 94% of messages are handled automatically, a person glances at 6%, and your inbox gets 16 spam messages out of 747.
Want to be even more careful with genuine enquiries? Moving the bands to 0.95 and 0.1 cut wrongly blocked genuine messages to just 1 and let only 7 spam through - but a person then checks 24% of messages. Your volume decides which trade is worth it.
What It "Got Wrong" - and Why That's Interesting#
We read the mistakes. Many weren't really Jev's fault.
Blocked genuine messages were mostly forwarded chain messages ("say this slowly…"), a clickbait celebrity link, and a job advert - all labelled "genuine" in the dataset, because a friend forwarded them. Most people would call them junk.
Missed spam included several jokes and one-liners labelled as spam in the dataset, which Jev scored as harmless. Again, arguable.
This is worth knowing for any AI test: your labels have errors too. When you measure accuracy on your own data, read the disagreements before you blame the model - some will be labelling mistakes, and some will reveal that your definition of "spam" needs to be clearer in the question.
Beyond Spam: Fit and Urgency#
Spam is the easy part of enquiry triage. What sales teams really want to know is is this a good lead and how quickly do we need to reply.
Jev can ask those questions on the same message, in the same request, in parallel:
- Fit - a score question with levels you define from your own sales criteria, for example "outside our services", "possible fit, needs a conversation", "clear fit: our core service, in our area, with a stated need".
- Urgency - a score question from "general research, no timeframe" to "needs help now or has a deadline this week".
We haven't published numbers for these, deliberately. There's no public dataset of labelled business enquiries, and a benchmark on made-up examples would tell you nothing. Fit depends entirely on your services and your customers, which is exactly why these need testing on your own past enquiries - ideally 200-500 of them, for the reasons in our confidence threshold test.
A caution: enquiries contain names, emails and phone numbers, and Jev is hosted in the US. Strip what you can before sending and read our UK safety guide first. Jev rarely needs someone's phone number to judge whether their message is spam.
Should You Filter Enquiries With Jev?#
It's a strong fit if you get enough form submissions for spam to waste real time, and you can route borderline scores to a person rather than deleting them.
Keep it simple if you get a handful of enquiries a week - a basic form filter and your own eyes are fine. See when not to use Jev.
Never auto-delete. Even at a 0.95 threshold, one genuine message in our test would have been blocked. Send high scores to a spam folder someone skims weekly, not the bin.
Frequently Asked Questions#
How accurate is Jev at detecting spam?#
On the 5,574-message SMS Spam Collection, Jev scored an AUC of 0.992. At a 0.9 threshold it caught 84.3% of spam with 98.6% precision, wrongly flagging 0.19% of genuine messages.
How much does spam filtering with Jev cost?#
About $0.015 per 1,000 messages in our test (roughly 360 input tokens each at $0.042 per million, output free).
What threshold should I use for spam?#
For contact forms, use two thresholds: send 0.9+ to a spam folder, 0.3-0.9 to a quick human check, and pass the rest. Then confirm the numbers on a few hundred of your own past submissions.
Can Jev score lead quality too?#
Yes - as separate score questions in the same request. But lead quality is specific to your business, so it must be tested on your own labelled enquiries before you rely on it.
Data: SMS Spam Collection by Tiago A. Almeida and José María Gómez Hidalgo, UCI Machine Learning Repository, licensed under CC BY 4.0. Tests run with jev-1.13.0.
Want to know how this works on your own enquiries or tickets? Our Jev Switch Check tests your real (anonymised) data against Jev and a backup model and tells you what you'd save and where it's safe. Start with a free savings estimate.
AI engineer at BrightBit Digital



