Most Jev tutorials stop at "hello world". This one is built from the code we used to send more than 11,000 requests to TypeSafe's Jev for our support routing and spam filter tests - including the parts that only matter once you run it for real: confidence bands, batching and retries.
By the end you'll have working Python for the two most useful question types, and you'll know how to turn Jev's answers into safe decisions in your product.
What You're Building Against#
Jev isn't a chatbot, so the API doesn't take a prompt and return text. Every request has three parts:
- state - the information Jev should look at (a message, a ticket, a document).
- questions - one or more typed questions about that state.
- model - which Jev version to use. We used
jev-latest.
Each question has a type:
| Type | Returns | Example |
|---|---|---|
noul | Probability a statement is true (0-1) | "Is this message spam?" |
choice | One option from your list, with probabilities | "Which team should handle this ticket?" |
score | A level on a scale you define | "How urgent is this?" |
We'll cover noul and choice - the two we've tested at scale. score works in a similar way; check TypeSafe's documentation for its exact format.
Step 1: Get a Key and Keep It Secret#
Create an API key in your TypeSafe account (Jev is in early access, so you may need to request access first).
Treat it like a password:
- Store it in an environment variable, never in your code.
- Only call Jev from your server. A key in browser JavaScript is a key anyone can copy.
- Make sure your
.envfile is in.gitignorebefore you add the key - and check git isn't already tracking it.
export JEV_API_KEY="your-key-here"
pip install httpx
Step 2: Your First Call - a Yes/No Question#
Here's a complete spam check. The noul question asks how likely a statement is to be true, and the criteria describe what true and false mean.
import os
import httpx
JEV_URL = "https://api.typesafe.ai/v1/systemone"
API_KEY = os.environ["JEV_API_KEY"]
def ask_jev(state: dict, questions: dict) -> dict:
response = httpx.post(
JEV_URL,
headers={"Authorization": f"Bearer {API_KEY}"},
json={"state": state, "model": "jev-latest", "questions": questions},
timeout=30,
)
response.raise_for_status()
return response.json()
SPAM_QUESTION = {
"spam": {
"type": "noul",
"instructions": (
"Is this message spam: unsolicited marketing, a scam, or prize or "
"offer bait, rather than a genuine message from someone the "
"recipient knows or deals with?"
),
"criteria": {
"true": "Spam, scam or unsolicited promotion",
"false": "A genuine personal or service message",
},
}
}
result = ask_jev({"message": "You've WON a £500 voucher! Reply CLAIM now"}, SPAM_QUESTION)
print(result["answers"]["spam"]["noul"]) # probability it's spam, 0 to 1
The question has an ID you choose ("spam"), and the answer comes back under the same ID in result["answers"]. That's how you match answers to questions when you ask several at once.
Step 3: Pick From a List - a Choice Question#
For routing, use choice. The criteria are your options: each key is what Jev returns, each value describes it.
ROUTING_QUESTION = {
"team": {
"type": "choice",
"instructions": "Which team should handle this customer message?",
"criteria": {
"billing": "Payments, invoices, refunds and charges",
"technical": "Bugs, errors, login problems, something not working",
"sales": "Pricing questions, upgrades, new orders",
"account": "Changing details, closing or verifying an account",
},
}
}
result = ask_jev({"customer_message": "I was charged twice this month"}, ROUTING_QUESTION)
answer = result["answers"]["team"]
print(answer["choice"], answer["confidence"])
Make the options clearer with examples#
This is the single biggest accuracy lever we found. Instead of a short description, give each option a few real example messages:
"criteria": {
"billing": {
"topic": "Payments, invoices, refunds and charges",
"example_messages": [
"Why was I charged twice?",
"Can I get a refund for last month?",
"Where do I download my invoice?",
],
},
# ...the same for each option
}
In our 77-topic test, three examples per option lifted accuracy from 80.8% to 88.8%. The trade-off is cost: the examples are sent with every request, so input tokens went from about 1,706 to 6,355 per call. At Jev's price that's still only about $0.27 per 1,000 calls - see our Jev pricing breakdown.
Step 4: Read the Response#
Here's a real response from our routing test (trimmed - the full probabilities list covers every option):
{
"model": "jev-1.13.0",
"answers": {
"topic": {
"choice": "card_arrival",
"confidence": 0.97,
"probabilities": {
"card_arrival": 0.98,
"card_delivery_estimate": 0.02
}
}
},
"usage": { "input_tokens": 6349, "output_tokens": 827 }
}
What each part is for:
choice- Jev's answer. Use this in your code.confidence- how sure it is. This is what you'll use to decide whether to trust the answer (step 5).probabilities- the full ranking. Handy for showing a person Jev's top three suggestions when it isn't sure. In our test, the right answer was in the top three 96% of the time.model- the exact version that answered.jev-latestcan change over time, so log this with every decision.usage- input tokens are what you pay for. Output tokens are currently free.
Step 5: Don't Trust Every Answer - Use Confidence Bands#
This is the step most tutorials skip, and it's the one that makes Jev safe to put in production. Never act on every answer. Split them by confidence:
def route(message: str) -> dict:
answer = ask_jev({"customer_message": message}, ROUTING_QUESTION)["answers"]["team"]
top_three = sorted(answer["probabilities"], key=answer["probabilities"].get, reverse=True)[:3]
if answer["confidence"] >= 0.9:
return {"action": "auto_route", "team": answer["choice"]}
if answer["confidence"] >= 0.5:
return {"action": "human_check", "suggestions": top_three}
return {"action": "manual_triage"}
The thresholds above are placeholders. A confidence of 0.95 doesn't mean "95% accurate" - in our test, answers around 0.95 were right 83-90% of the time. Set your thresholds on your own labelled data. Our confidence threshold guide shows how many examples you need (short answer: 200-500).
Step 6: Run Thousands of Requests Safely#
For batches - backfilling old tickets, or testing on your history - send requests concurrently, but not all at once, and retry temporary errors. This is the pattern we used for all our tests: eight requests at a time, retrying rate limits and server errors with a growing delay.
import asyncio
import random
RETRY_STATUSES = {429, 500, 502, 503, 529}
async def ask_jev_async(client, state, questions, retries=6):
for attempt in range(retries):
try:
response = await client.post(
JEV_URL,
headers={"Authorization": f"Bearer {API_KEY}"},
json={"state": state, "model": "jev-latest", "questions": questions},
timeout=60,
)
except httpx.HTTPError:
response = None
if response is not None and response.status_code not in RETRY_STATUSES:
response.raise_for_status()
return response.json()
await asyncio.sleep(2 ** attempt + random.random()) # 1s, 2s, 4s... plus jitter
raise RuntimeError("Jev request failed after retries")
async def classify_all(messages, concurrency=8):
limit = asyncio.Semaphore(concurrency)
async with httpx.AsyncClient() as client:
async def one(message):
async with limit:
return await ask_jev_async(client, {"customer_message": message}, ROUTING_QUESTION)
return await asyncio.gather(*(one(m) for m in messages))
results = asyncio.run(classify_all(["I was charged twice", "The app won't log me in"]))
Run this way, all of our 11,000+ calls succeeded, with a typical response time of about a quarter of a second each.
Mistakes to Avoid#
- Asking two things in one question. "Is this urgent and a good lead?" is two questions. Ask them separately - you can include several questions in the same request.
- Sending the whole record. Irrelevant fields distract the model and cost tokens. Send only what the question needs. It also keeps personal data out of the request.
- Doing maths or dates with Jev. Work out "is this invoice overdue?" in code and pass the result in. More in when not to use Jev.
- No fallback. TypeSafe's current terms have no uptime guarantee. If Jev is down or slow, your code should fall back to another model or a person, not crash.
- Forgetting UK data rules. Jev is hosted in the US. Read our UK safety guide before sending personal data.
Frequently Asked Questions#
What is the Jev API endpoint?#
We sent requests to https://api.typesafe.ai/v1/systemone with the API key as a Bearer token, and jev-latest as the model. Check TypeSafe's documentation for the current details, as the product is in early access.
Is there a Python SDK for Jev?#
You don't need one - it's a single JSON POST request, so any HTTP client works. Frameworks are adding support too; Pydantic AI, for example, documents TypeSafe as a model provider.
What's the difference between noul and choice questions?#
A noul question returns the probability that one statement is true - good for yes/no checks like spam. A choice question picks one option from a list you provide and returns probabilities for each - good for routing and categorising.
How fast is the Jev API?#
In our tests from the UK, the typical response took about a quarter of a second (248-273 ms median), including the round trip to TypeSafe's US servers.
Can I ask Jev several questions in one request?#
Yes - the questions field takes several questions, each with its own ID, and the answers come back under the same IDs. Keep each question to a single judgment.
Rather not wire it up yourself? Our Jev Switch Check finds the AI calls in your product that suit Jev, tests them on your real data, tunes the confidence bands, and hands over drop-in code. Start with a free savings estimate.
AI engineer at BrightBit Digital



