Dev.to AI 🤖 Ai 👁 0 📖 8 min read

No Reference Class: The Decisions Your Past Has Never Made

No Reference Class: The Decisions Your Past Has Never Made An underwriter who has quoted a thousand cargo policies will quote the thousand-and-first before lunch. Ask the same underwriter to quote the first hull built

No Reference Class: The Decisions Your Past Has Never Made

An underwriter who has quoted a thousand cargo policies will quote the thousand-and-first before lunch. Ask the same underwriter to quote the first hull built to a new design, with no voyages behind the design, no claims record, and no sister ship anywhere in the book, and something changes. They go quiet. They ask for time. Nobody in that room would say the second question is harder to compute. It is harder because there is nothing behind it to lean on.

That pause is the subject of this piece, because it is the one thing a judge that answers in milliseconds cannot be improved into doing. Not because it is slow, and not because it is costly. Because the material it is made of runs out.

Two different reasons a judge says it does not know

Take any model that returns a choice with a confidence value attached, ours included, and separate the ways it can be unsure.

The first is statistical. The case is ordinary, the precedents point in different directions, both candidates survive a careful reading, and the reported number lands near the middle. You have seen this shape before and so has the model. The reference class exists and it is crowded.

The second is structural. The case has no precedent at all. Not a contested one — none. Your company has never made this call, the industry has no settled answer, and the closest thing in the training data is a different kind of decision that happens to share vocabulary.

On a dashboard these look identical. A low number is a low number. They are not the same finding, and only one of them can be fixed by collecting more data.

A confidence value is a claim about a population

Read the number slowly. When a system reports 0.9, it is claiming that among cases that look like this one, roughly nine in ten turn out the way it says. The phrase doing all the work is cases that look like this one. A probability is not a property of a question. It is a lookup into a set of comparable questions that have already been answered.

When that set is large, the number is meaningful and testable: count the outcomes and compare. That is what a reliability table is for, and it is why a vendor that will not publish one has not measured anything. When the set is empty, the number is still printed, still formatted to two decimals, and still a number. It has simply stopped referring to anything. What it inherits is the shape of the nearest population the model has seen — a different question, with a different base rate, and consequences that land somewhere else.

Be precise about how this differs from the more familiar complaint about distributions. That complaint says the population you are running on is not the population the number was measured on: the mix is different, the base rate has moved, and the curve is stale. It is a serious problem and it is fixable in principle, by re-measuring on your own cases. The case here is not that. It is not a mismatched population, which can be measured again. It is an absent one, which cannot be measured at all.

So the boundary is not "the model is not smart enough yet." A larger model, trained on ten times as much of your history, still has no answer for a case that has never occurred, because the answer is not in the history. It has not been made yet.

Where the reference class exists, we can be held to it

Everything above would be cheap rhetoric if the people making the argument refused to be measured on the part that is measurable.

Ours is self-run under named benchmarks, with failures disclosed rather than quietly retried. On JudgeBench: 620 judgments, with the 6 first-verdict failures disclosed. Raw accuracy came out at 92.5% against 92.2% for a plain direct baseline. That is a tie, and we report it as a tie; the claim here is not a sharper judge. The bands are the part worth keeping: calls we reported at 90% confidence or above came back right 99.6% of the time, and the 80–90% band 94.0%. On ContextualJudgeBench, self-run under the official pairwise protocol where a pair counts as correct only if both presentation orders are judged correctly — which is why the random floor is 25% rather than 50% — consistent accuracy is 67.1% against the benchmark's official reference of 65.4%, with 12 orders excluded after repeated platform failures and that exclusion stated next to the result. The deliberately constructed near-tie splits sit at 46–60%.

Put that last figure where this argument needs it. A near-tie is a case whose comparable set is internally contradictory: the population exists, and it points both ways at once. At 46–60% the instrument is reporting that the reference class has dissolved into noise. That is not a defect to be hidden. It is the only honest signal a frequency-based judge can send at the moment frequency stops carrying information — and it is why the useful output there is not a better number but a person.

The fast end got its benchmark within a week

There is a telling asymmetry in how quickly the two kinds of judgment acquired measuring instruments.

Within days of Jev appearing, the ecosystem was building comparisons for the precedent-bound part. "Show HN: JevBench, a reproducible benchmark for typed decision models" reached 154 points on Hacker News on September 22, and "Jeeves. Reasoning improves Jev-like decision models" reached 242 points on September 29. Those are per-post scores read off the public search index, and they drift by a few points over time; the pattern does not. Where the question is "which of these labels fits," a corpus is free, the labels are already attached, and a benchmark can be assembled in an afternoon by someone with no access to the original model.

For unprecedented decisions there is no equivalent, and the absence is not laziness. You cannot assemble a corpus of decisions that have not been made. You cannot hold a benchmark constant when the ground truth is the thing being selected. The fast end has a leaderboard because it has a yardstick. The slow end has neither, which is why the slow end is still argued about in prose.

Unfamiliar is not the same as unprecedented

The mistake in the other direction is more common, and it burns more minutes than the first one.

Most decisions that feel novel are merely unfamiliar to the person holding them. The pricing question that has never come up inside your company has been answered a thousand times in your industry. The build-or-buy call that feels unprecedented is in somebody's public post-mortem. Novelty that exists in your head but not in the world is a millisecond decision wearing a costume, and treating it as a minute decision is how a team spends a quarter rediscovering a base rate that was already published.

So the discipline has to run in both directions, and it comes down to three questions.

Does a comparable case exist anywhere — your own history, a competitor's, a dataset, a public record? If it does, this is a lookup, and a lookup should be done cheaply and at volume.

Would more data change the answer? If the answer is a fact about the world, then data helps and you should go get it. If the answer depends on what you want to become — which of two futures you are willing to own — then no additional data settles it, because the data describes the world you have, not the one you are choosing.

Will somebody cite this decision later? If the next person facing this exact call will point at what you did as their precedent, then you are not looking up an answer. You are writing the first entry in a record that will be read by people who were not in the room.

What you can hand to a decision that has no precedent

The second kind of decision needs a different artifact from the first kind.

It does not need a probability, because there is no population to be right about. It needs to be written as a first case rather than a verdict: the reasoning spelled out in sentences that a stranger could carry to the next instance of the same question and land in the same place. It needs both candidates laid out, not only the one you already prefer. It needs the argument written down, because the written argument is the part that can be disagreed with, corrected, and cited by the next person who arrives here. And it needs a name, because the first entry in a record is also the moment liability attaches.

That is the honest division of labor. The cheap end is a lookup, and lookups should be fast, numerous, and boring. The expensive end is a precedent being set, and a precedent is not retrieved — it is authored, once, by someone who can be asked about it later.

We build for the second case, so it is fair to state plainly what we can and cannot put on the table. Where a reference class exists, we publish bands with counts and with exclusions, which is how you can tell whether a number we report is worth acting on. Where one does not exist, we do not have a calibrated probability to offer, and we will not dress one up. What we can produce is a pick between two candidate answers, a confidence stated together with the population it was measured against, and a written argument long enough to disagree with — which, for a decision whose entire difficulty is that no precedent covers it, is the artifact you actually need.

If the call in front of you is unprecedented rather than merely hard — one question, two defensible answers, and a record that will be cited — that is the case this exists for: Decider.

📰 Read the original article on Dev.to AI

Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.