What is Jev? The AI trained to say “I'm only 60% sure.”
I kept hearing the word. Jev. In group chats, in newsletters, in a Hacker News thread that hit 1,900 points and 500 comments in a day — which for that site is a small riot. Every explanation I clicked on was written for
I kept hearing the word. Jev. In group chats, in newsletters, in a Hacker News thread that hit 1,900 points and 500 comments in a day — which for that site is a small riot.
Every explanation I clicked on was written for engineers. Non-autoregressive. Calibrated posteriors. Typed schemas. I understood maybe half.
So I read the launch post, the docs, and all 500 comments. Underneath the jargon is a simple idea, and a slightly funny one.
Every AI you've used so far writes. Jev doesn't. Jev decides.
Stick with me. By the end you'll be able to explain it to a friend in one sentence, you'll know what the “can't hallucinate” claim actually means (not what it sounds like), and you'll know whether it ever touches your life. It will. You won't see it.
The one-sentence version
Think about the difference between an essay and a light switch.
ChatGPT, Claude, Gemini — they write essays. You ask, they produce words, one after another, and the words can be anything. A poem. A recipe. Working code. A confident lie. That flexibility is the magic and the problem.
Jev is a light switch. You don't ask it to talk. You show it a situation and ask a fixed question with a fixed set of answers — is this urgent, yes or no? which of these five teams handles it? how angry is this customer, one to five? — and it flips the switch. Instantly. And next to the switch it puts a number: how sure it is.
It cannot write you a sentence. That isn't a limitation they're fixing; it's the design. The people who built it gave up words on purpose, because they think most of what businesses want from AI isn't words. It's decisions.
Why it's called that
Two names, both borrowed from smart dead people, and both help.
“System One” is Daniel Kahneman's. He split thinking into two modes. System 2 is slow and careful — doing your taxes. System 1 is fast and automatic — you see a face and know it's angry before you could explain why. You don't reason your way there. You just know, and you're usually right.
Chatbots are being pushed towards System 2 — “think step by step”, reasoning modes, minutes of pondering. Jev is built for System 1: the thousand tiny snap judgements that happen inside software all day. Is this spam? Is there a person in this photo? Which folder does this go in? Nobody wants an essay about it. They want the answer, fast, and they want to know if it's reliable.
“Jev” is after William Stanley Jevons, the economist who noticed that when steam engines got more efficient, people didn't use less coal. They used vastly more, because cheap power made a thousand new things worth doing. The founders are betting the same happens when a decision costs almost nothing. That's the whole business plan, in a name.
The founder is Diogo Almeida, who was at OpenAI on the research that taught language models to follow instructions — the work that became ChatGPT. So: someone who helped build the essay machine, deciding the next thing shouldn't be one. The company, TypeSafe AI, reportedly raised $40 million to do it.
What you actually send it
A real example from their documentation. It's the moment it clicked for me.
You give Jev a situation — they call it “state”. Say, a support message:
“Hi, I've been trying to connect my Stripe account for 3 days and it keeps failing. I'm losing sales. Please help ASAP.”
And a question with a fixed shape. Here, a yes/no: does this message convey urgency?
What comes back is not a paragraph. It's this:
is_urgent: 0.999
A number. 99.9% yes. Your software reads it and acts — bumps the ticket, pings a human, whatever you've set up. No “Certainly! This message appears to be…” to wade through.
There are only three kinds of question you can ask:
- Yes or no. Returns the probability of yes.
- Pick one from a list. Returns a probability for every option, plus an overall confidence. Up to 255 options.
- Score it on a scale. Low/medium/high, one to ten, whatever you define. Returns the score, the spread, and the confidence.
The trick that makes it fast: you can ask dozens of these about the same situation at once, and it answers all of them in a single pass. A chatbot builds its answer one word at a time, each depending on the last. Jev builds nothing. It looks once and flips every switch at the same time.
Here's the difference as a picture:

A chatbot answers word by word. Jev answers every question at once.
That's why the speed numbers are wild. They claim 70 to 500 milliseconds per answer against the three seconds to five minutes a reasoning chatbot takes. And pricing: $0.042 per million words of input, output free, because the output is a handful of numbers. Someone on the team built a bot that plays Doom by asking Jev ten questions a second. About $7 an hour. The team's reaction was that this was cheaper than they expected.
Honest note: those are TypeSafe's own numbers, from their own laptops, on tasks they chose. I haven't run it — it's early access with a waitlist. More on what independent people made of the claims below.
The claim everyone repeats
Every headline says the same thing: Jev can't hallucinate.
This is true. It's also not what you think it means, and the gap between those two is the most useful thing in this post.
A chatbot “hallucinates” when it confidently produces something false — a fake court case, a made-up statistic, a function that doesn't exist. It can, because it can write anything. The space of possible outputs is infinite, and some of that space is nonsense.
Jev can't write anything. You gave it a menu. Yes or no. One of these five teams. A number from one to ten. It is mathematically impossible for it to hand you something that isn't on the menu. No fake court case, because “fake court case” was never an option.
But it can absolutely pick the wrong thing off the menu.
If the message wasn't urgent and Jev says 0.9 urgent, that's a wrong answer. Not a hallucination — a mistake. Several of the sharpest people in that thread made exactly this point: it can't emit an invalid answer, but it can still emit a completely wrong valid one.
So why is it still a big deal? Because of the number. Here's the honest version on one card:

The new part isn't the left card. It's the number on the right — trained to be honest.
What TypeSafe actually built — the new part — is that Jev is trained to make that confidence number honest. Their training method is literally called Reinforcement Learning for Calibrated Decisions. Calibrated means: when it says 90%, it should be right about nine times in ten. When it says 55%, it should be a coin flip, and it should say so.
Their own post puts the problem perfectly: “If a model can do a task 95% of the time but doesn't say when it's in the 5%, it can't automate that task.” That's the whole reason AI has been stuck in “assistant” mode instead of “just do it” mode. Not that it's wrong 5% of the time — that you can't tell which 5%.
There's good research showing chatbots make people more confident and less accurate, because they never say “I don't know”. Read that sentence again with Jev in mind. This is an AI whose entire training objective is to say how unsure it is. Whether it lives up to that, nobody outside the company has tested yet. But it's the right thing to be trying to build.
What the sceptics said, in plain words
Five hundred comments. The best ones, translated.
“This already existed.” True-ish. Machine learning has had classifiers for years — models that take input and output a probability, fast, no words. Your spam filter is one. What's new, several people concluded, is that you don't have to train one per job. You describe the menu in plain English and it works immediately. One commenter called it “democratisation of classifiers”. Fair, and still a big deal — training a classifier used to need an ML engineer and a pile of labelled data.
“The speed comparison is apples to oranges.” Also fair. Jev is 200× faster than a chatbot at flipping switches. A chatbot can also write code, draft your email, and explain the switch. Comparing them on speed is like saying a light switch is faster than a novelist.
“Their evals grade against other AIs, not against truth.” This one matters. TypeSafe's benchmark checks how closely Jev's decisions match the average of GPT-6 and Claude's flagship on the same tasks — not how often it is actually right. So “as good as the best models” really means “agrees with the best models”. Those models can be wrong together.
“It's probably a small model, and people will copy it.” One commenter estimated from the price that Jev is around 3 billion parameters — tiny by today's standards — and predicted clones within weeks. Three days after launch, “Cua S1 — a family of System One models” showed up on the same site. The category may matter more than the company.
Credit where it's due: TypeSafe's launch post has a “Nuance” section under every claim, admitting where their numbers flatter them, where bias could exist, and that they can't prove their pricing isn't subsidised. More honesty than most launches manage.
Where you'll actually meet it
You will never open an app called Jev and type at it. That's the point. It's plumbing. It shows up inside things you already use, and the tell will be that AI-powered features get faster and quieter. Here's where it would sit:

Read the last box. The uncertain one goes to a person. That's the pattern this makes possible.
That last box is the bit that affects you. Route to a human when confidence is low. The AI handles the 90% it's sure about; the uncertain 10% goes to a person. Today most AI features either handle everything (and get some wrong, confidently) or nothing.
If it works as claimed, the things that get better are boring. Support tickets to the right person first time. Spam and fraud checks that run in a blink. Game characters that react instead of freezing to think. Photo apps that sort ten thousand images without a queue. Nothing you'd write a headline about. Everything you'd notice if it stopped.
Paragraph or checkbox?
The practical part, and it's for anyone — not just coders. Plenty of people automate things with no-code tools now, and all of them are about to face this choice.
When you're thinking of using AI for a task, ask one question first: is the output a paragraph or a checkbox?
If a human needs to read the answer — a draft, an explanation, a summary — you want a chatbot. If software needs to act on the answer — sort, route, flag, score, approve — you want a decision model, and Jev is the first mainstream one. Here's the whole decision:

Paragraph or checkbox? Which kind of AI your task actually needs.
The smartest comment in the whole thread was about the box in the middle. It's not Jev versus ChatGPT. The likely pattern is both: you use a chatbot to design the questions — what should the menu be? what does “urgent” mean for us? — and the decision model runs those questions a million times in production. The essay-writer designs the switchboard. The switch-flipper runs it.
Who should ignore this
If you use AI to write, learn or think — most people — Jev changes nothing for you today. Carry on.
If you build anything, automate anything, or decide how AI gets used where you work, it's worth twenty minutes on the docs and the waitlist, with expectations set: early access, self-reported numbers, a menu-bound model that can still be wrong, and a category that will have five competitors by Christmas.
And if you just like knowing where this is going — the interesting bit isn't the speed. It's that someone who helped build the most confident machine in history just built one whose job is to admit doubt.
The sentence to tell your friend
“It's an AI that doesn't write anything — you give it a situation and a multiple-choice question, and it picks an answer instantly and tells you how sure it is.”
If they say “so it's a spam filter”, say “yes, but you can build one in a sentence instead of a month.” If they say “so it can't be wrong”, say “it can — it just can't be impossible, and it tells you when it's guessing.”
Got early access and pointed it at something real? Tell me — especially whether the confidence numbers meant what they claimed. That's the one thing nobody outside the company has answered yet.
This is the thing we build. AI that says “I don't know” instead of guessing, and hands the uncertain ones to a person — inside the tools your business already uses. Jev is one way to get there; a well-built assistant over your own material is another. Either way the rule is the same: the number next to the answer has to be honest.
Sources: TypeSafe's launch post and docs (the Stripe example is verbatim from their quickstart) and the Hacker News thread of 15 Sep 2026. Speed, price, the $7/hour figure and the 255-option limit are TypeSafe's stated numbers. I have not used Jev.
Read next: You gave AI your documents. It’s still wrong.
Originally published at singhlabs.dev.
Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.