Dev.to AI πŸ€– Ai πŸ‘ 0 πŸ“– 8 min read

Jev vs small LLMs: does a model that decides beat one that writes?

Key findings Jev tied two small LLMs on accuracy, beat them on probability quality and speed but failed completely at weekday arithmetic. Accuracy is a tie. On 150 CLINC questions Jev scored 0.980, gpt-oss-20b 0.963

Jev vs small LLMs: does a model that decides beat one that writes?

Key findings

Jev tied two small LLMs on accuracy, beat them on probability quality and speed but failed completely at weekday arithmetic.

  • Accuracy is a tie. On 150 CLINC questions Jev scored 0.980, gpt-oss-20b 0.963 and gpt-oss-120b 0.923. All three 95% intervals (the range a score could plausibly land in on a different batch of questions) overlap, so this sample cannot rank them.
  • Jev's probabilities are better. Its Brier score (how far the probabilities it gave were from the truth, lower is better) is 0.037 against 0.071 for the 20b, about half the error. Its NLL (a penalty that punishes confident wrong answers really hard, also lower is better) is 0.064 against 0.237.
  • Jev is 7-8x faster. Median latency is 0.40 s against 2.76 s and 3.24 s even when using Groq's services whose whole schtick is faster inference. The whole Jev experiment cost about one cent ($0.0104).
  • Jev fails on weekday questions. It scored 0.13 where random guessing scores 0.14. The 120b got every blind-spot question right, the 20b 0.89, Jev 0.61.
  • The 120b's CLINC numbers are partly a harness artifact. 15 of its 300 replies failed to parse (couldn't be read as the JSON it was asked for) and were scored as random guesses. Fix this before publishing.

What's Jev?

Jev made huge waves in the ai community, but what exactly is it?
Well it's a model that decides instead of writing: you give it text and a list of options, and it returns a choice plus a probability for every option, that's it!

Any normal LLM is generative in nature and would answer the same question with a sentence, which your code then has to read and parse and whatnot. Jev's answer is already a typed value your code can use. TypeSafe AI, its creator, calls this class System One models.
The name's borrowed from Daniel Kahneman's philosophy, where System 1 does fast intuitive judgement whereas System 2 does slow deliberate reasoning.

Jev was released on 15 September 2026. On OpenRouter(that's how I got access to this model) it is typesafe/jev-1.13, priced at $0.042 per million input tokens with free output(It's literally so cheap that it's free because there's a set limit to how much it can output/choose).

While I was reading about this model, a thought came to my mind: can a specialist that only decides beat small general-purpose LLMs at deciding? For this I conducted a very simple experiment. Speed and cost are Jev's selling points, so I also measured whether the answers hold up.

The experiment

Every model answered the same multiple-choice questions, each with 16 options (7 for the weekday questions), twice with the options shuffled differently to see if it makes a difference.

Model How it was called Settings
Jev 1.13 (typesafe/jev-1.13) OpenRouter Decisions API Native probabilities & no temperature setting
gpt-oss-20b Groq, free tier Temperature 0, low reasoning effort, 400-token reply cap
gpt-oss-120b Groq, free tier Same as the 20b

Reasoning effort basically controls how much hidden thinking these models can do before answering. Temperature at 0 makes it as deterministic as possible.

Three question sets

  1. CLINC-150 (main). 150 customer-assistant requests from the public CLINC-150 test split with out-of-scope rows removed. The model is given one 'true intent' option with 15 other incorrect ones.
  2. Blind-spot set (50 questions I built to stress it). 20 arithmetic (3-digit by 2-digit multiplication), 15 letter counts (how many "s" in "mississippi"), and 15 weekday questions ("what weekday is it N days after this date", with all 7 weekdays as options).
  3. Variance set. The first 20 CLINC questions, sent to Jev three identical times each, to see whether it answers the same way twice.

How each model was asked. Jev received the options as its native choice format and returns probabilities directly. The LLMs got lettered options and were told to reply with JSON probabilities summing to 1. A reply that could not be read counted as a parse failure and was scored as an even guess(basically all options are given the same probability. Since we pick the one with the highest one, we pick the first option that's seen. There's basically a 1 in 16 chance the first option is correct). All models saw identical option sets and identical orderings, from a fixed random seed.

Result 1: accuracy is a statistical tie

Experiment log Β· CLINC-150, 150 questions x 2 orderings

Model Accuracy 95% interval Parse failures
Jev 1.13 0.980 0.953 to 1.000 0 of 300
gpt-oss-20b 0.963 0.933 to 0.990 0 of 300
gpt-oss-120b 0.923 0.883 to 0.957 15 of 300

Jev's 0.980 is the highest score, but its 95% range overlaps both LLMs'. With 150 questions, gaps of a few points are basically noise, so this sample can't truly suggest which model is more accurate.

The 120b's lower score is largely caused by it's harness. Its 15 unreadable replies count as guesses. Scoring only its 285 readable replies gives roughly 0.97, in line with the other two. That is a back-of-envelope estimate, not a re-run.

Result 2: Jev's probabilities are better

Experiment log Β· CLINC-150, 300 calls per model Β· Brier and NLL over all 16 options, ECE with 5 bins Β· no confidence intervals were computed for these metrics

Model Brier (0 best, 2 worst) NLL (lower is better) ECE, 5 bins (0 best)
Jev 1.13 0.037 0.064 0.013
gpt-oss-20b 0.071 0.237 0.096
gpt-oss-120b 0.096 0.288 0.052

Brier : Checks how far probabilities were from truth.
NLL(Negative Log Likelihood) : Looks at probability of put on correct answer, and sees the negative log of it. For example, 0.9 on correct answer gives 0.11
ECE(Expected Calibration Error) : Checks if the confidence matches truth. Groups calls into 5 bins based on how confident model was and then compare average confidence with actual accuracy in bin.

Against the 20b, Jev's Brier score is about half (0.037 against 0.071), its NLL is 3.7 times lower, and its ECE is about 7 times lower.

The 120b looks worst, but 15 parse failures are scored as even guesses, each worth a Brier of 0.94. Setting them aside, its Brier drops to about 0.052 and its NLL to about 0.157. That beats the 20b and still trails Jev. These are estimates worked back from the recorded totals, not a re-run.

Two cautions. Only accuracy got a confidence interval, so treat the size of these gaps as approximate; the direction is consistent across all three metrics against the 20b, which had no parse failures. And the LLMs wrote their probabilities as text, while Jev emits them natively, which handicaps the LLMs.

Result 3: Jev is 7 to 8 times faster and cost about a cent

Experiment log Β· CLINC-150, 300 calls per model, run 30 Sep 2026 Β· times include network round trip

Model Median (p50), s Slowest 5% (p95), s p50 vs Jev p95 vs Jev
Jev 1.13 0.40 0.74 1.0x 1.0x
gpt-oss-20b 2.76 3.85 6.9x 5.2x
gpt-oss-120b 3.24 4.72 8.1x 6.4x

Jev's median call took about 0.40 s against 2.76 s for the 20b and 3.24 s for the 120b, so 6.9 and 8.1 times faster. The slow tail is far shorter too. At p95 (the time 95% of calls beat, so basically the slowest ones) Jev took 0.74 s against 3.85 s and 4.72 s.

The entire Jev experiment cost about $0.0104: $0.0069 for the 300 main calls, $0.0022 for the blind-spot set and $0.0014 for the variance run. That works out to about 2.3 cents per thousand calls at this prompt size. The LLM calls ran on Groq's free tier.

One caution: these times include the network. Jev was reached through OpenRouter and the LLMs through Groq, so this compares two services as I used them, not the models' raw speed.

Result 4: Jev's answers did not depend on option order

On CLINC, shuffling the options never changed Jev's top answer, and repeating identical calls never changed it either. Flip rate is the share of questions whose top answer changed between the two orderings.

Question set Jev gpt-oss-20b gpt-oss-120b
CLINC (150 questions) 0.000 0.007 0.087
Blind-spot set (50 questions) 0.04 0.12 ~0.000

Jev also answered identically when I sent the same 20 CLINC questions three times each (60 calls, flip rate 0.000). An independent study cited on the JEV-9B model card found Jev changes its answer on 4.3% of repeated 64-option calls, so my 16-option, 20-question check does not rule out occasional wobble.

The 120b's 0.087 is 13 of 150 questions, close to its 15 parse failures. An unreadable reply becomes an even guess whose top pick simply follows the option order, so those failures probably cause most of its flips. I have not confirmed this. The check is to recompute the flip rate with failed calls excluded.

On the blind-spot set the counts are tiny: 2 of 50 questions flipped for Jev and 6 of 50 for the 20b.

Where Jev starts showing it's weakness: weekday questions

Experiment log Β· 50 hand-built questions (20 arithmetic, 15 letter count, 15 weekday), 2 orderings, gold answers computed, not written by hand

Measure Jev 1.13 gpt-oss-20b gpt-oss-120b
Overall accuracy (100 calls) 0.61 0.89 1.00
Arithmetic (20 items) 0.95 0.97 1.00
Letter count (15 items) 0.63 0.73 1.00
Weekday (15 items; chance = 0.14) 0.13 0.93 1.00
Brier 0.422 0.205 0.000
Flip rate 0.04 0.12 0.00

On the 50 hand-built questions Jev scored 0.61, against 0.89 for the 20b and 1.00 for the 120b. Nearly all of that gap is weekdays. Jev got 0.13 right, and guessing among 7 weekdays scores 0.14.

A plausible reason, which I did not test: working out a weekday means a small calendar calculation, and Jev returns no reasoning trace while the LLMs think before they answer.

Two cautions. Each category has only 15 to 20 questions (30 to 40 calls), so read the bars as warning signs, not measurements. And the arithmetic questions are multiple choice, with wrong answers within 400 of the true product. That is far easier than computing a product from scratch, so Jev's 0.95 does not show it can do arithmetic.

What this does and doesn't show

This was a small experiment, so here's how far I'd trust it:

  • Unequal confidence. The LLMs had to write their probabilities out as text while Jev gives its own natively, so the LLMs were a bit handicapped on Brier, NLL and ECE.
  • The 120b's parse failures. 15 of its 300 replies (5%) couldn't be read and got counted as even guesses, which drags down its accuracy, Brier, NLL and probably its flip rate. I didn't check why. My guess is it hit the 400-token reply cap, but that's just a hypothesis.
  • Small samples, one run. About 150 questions, one seed, one run per condition. Gaps of a few points are basically noise, and Jev's and the 20b's accuracy ranges overlap. Only accuracy got confidence intervals.
  • Closed-set questions. The right answer is always somewhere in the 16 options. Real classification doesn't hand you that guarantee.
  • Untuned LLMs. I only tried reasoning_effort=low on Groq, and Groq doesn't report a model version, so I couldn't pin it. Jev was served as the dated snapshot typesafe/jev-1.13-20260917.
  • A small, artificial blind-spot set. 15 to 20 questions per category, with hand-made option sets.

So, is Jev worth it?

On this small test, Jev looks like a great fit for fast, fixed-option decisions where you need confidence you can actually trust, and a poor fit for anything that needs a calculation.

It matched the LLMs on accuracy, gave better calibrated probabilities than the 20b, and answered way faster for about a cent, even against Groq, whose whole schtick is fast inference. But it sat at chance on weekday questions, so anywhere the answer has to be worked out, I'd put an LLM or plain code behind it.

Sources

πŸ“° Read the original article on Dev.to AI

Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes β€” full credit and traffic to the original publisher.