Dev.to AI ๐Ÿค– Ai ๐Ÿ‘ 0 ๐Ÿ“– 6 min read

How I calibrated an LLM judge to grade like me, 25 cheaper

Businesses that sell technical products answer the same kind of question every day. "What's the accuracy on this range?" "Can it measure through a coating, and how thick?" "Does it come with a calibration certificate?" T

How I calibrated an LLM judge to grade like me, 25 cheaper

Businesses that sell technical products answer the same kind of question every
day. "What's the accuracy on this range?" "Can it measure through a coating, and
how thick?"
"Does it come with a calibration certificate?" The answers sit in
product datasheets and manuals.

I took 27 of those PDFs, three of them scans with no text layer, and wrote 47
questions the way customers actually ask them. Then I had four AI systems answer
every question and graded all 188 answers.

Running the systems took an afternoon. Building a judge I could trust took the
rest of the work, and that is what this post is about.

The result

Correct answers out of 47: my pipeline 46, Claude app 45, ChatGPT app 35, Chatbase 33

System Correct Wrong Made up
My pipeline 46 1 0
Claude app 45 1 0
ChatGPT app 35 10 0
Chatbase 33 11 0

My pipeline is Claude Sonnet over the API with page citations. The Claude app ran
Opus with the documents in a project, ChatGPT ran with thinking off, and Chatbase
was on the free plan with its default model.

Every system got the same PDFs and the same instructions. One run each, so I read
a one-question gap as a tie.

No system invented a spec value. I checked every claim twice: a claim checker
read each one against the page it came from, and I went through the ones it was
unsure about by hand. The weaker systems failed more quietly. They said an answer
wasn't in the documents when it was.

Where the misses came from

Correct answers by question type for each system

Scanned pages. Systems that only read the text layer saw blank pages. Every
question answered only by a scan came back "not specified".

Answers that need a second look. A product advertises one range on the front
page. A different mode of the same product stops at a tenth of it, and the
customer's question is about that mode. The same pattern showed up in accuracy
tables split by range and in specs that change with the material.

Unit conversions. The customer asks in one unit. The datasheet lists the
other unit in one row and a separate, lower limit in the customer's unit in
another. Converting the first number gives a confident wrong answer.

The website and the datasheet disagreeing. The product page said one value
and the datasheet said half of it. My rule: the datasheet wins, and the reply says
the page may be wrong. Two systems missed it.

Rules that live in no PDF

Before running anything I reviewed every question. I dropped three that no
customer would ask, corrected two gold answers, and turned the "not in the
documents" cases into rules:

  • Datasheet beats product page. Give the datasheet value and say the page may have an error.
  • Never a bare "not in our documents". Say it isn't in the published documents, and give an email to ask.
  • Certifications only if the document says so. Otherwise, email for confirmation.
  • Canned answers for the two most common questions.

After grading I added one more: answer what was asked. Extra conditions
confuse buyers and invite more questions.

That short list did more for answer quality than any prompt trick.

Why the judge needs a judge

Grading 188 answers by hand takes hours, so the usual move is an LLM judge: give
a model the question, the gold answer and the answer, and ask for correct, partial
or wrong. You only know the judge is any good if it agrees with the person whose
standard matters. Here, that's me.

So I graded 20 answers blind: no system names, five per system, with some likely
mistakes mixed in so the judge would be tested on errors too. Then I compared.

First try: 13 of 20. Most of the gap was in my own answer key. Three gold
answers were wrong. Grading real answers showed me that I answer from the
datasheet as printed, even when a conversion on it looks off, and flag the
document separately. My key hadn't been written that way. Once I fixed it,
agreement jumped.

Plain agreement flatters a judge when most answers are correct. A judge that
marks everything "correct" would agree 80% of the time here and understand
nothing. Cohen's kappa corrects for that.

Cohen's kappa: 67 of 100 agreement expected by luck, 28 from skill, 5 missed; kappa = 28 / 33 = 0.85

Kappa asks how far above luck the judge got, as a share of how far above luck it
could have got. 0 means no better than random; 1 means it matched me every time.

Two judges

Agreement with my grades: Sonnet 35 of 40, hybrid 36 of 40. Cost for 188 answers: $0.80 vs $0.03

Judge Agrees Cost, 188 Speed
Sonnet alone 35 / 40 ~$0.80 seconds
Hybrid 36 / 40 ~$0.03 under 1 s

The hybrid gives each part the job it does best.

The hybrid judge: code compares numbers, Jev answers typed questions, code applies the policy; Sonnet with the PDF only for unsure claims

Code pulls the numbers out of the gold answer and the answer being graded,
and records which match. It never confuses ยฑ(1.2% + 5) with ยฑ(0.8% + 5).

Jev handles meaning. It's a System One model from
TypeSafe: it reads natural language and returns typed answers with
probabilities, with no prose to parse. One request per answer asks four
questions at once:

{
  "verdict": {"type": "choice",
    "instructions": "Grade `answer` to `customer_question` against `gold` and `notes`. Use `code_checks` for which gold numbers the answer contains.",
    "criteria": {"correct": "...", "partial": "...", "wrong": "..."}},
  "brush_off": {"type": "noul",
    "instructions": "Does `answer` brush the customer off, such as a bare 'not in our documents' with no next step?"},
  "unasked_extra": {"type": "noul",
    "instructions": "Does `answer` add specs, conditions or other models that `customer_question` didn't ask about?"},
  "rule_followed": {"type": "noul",
    "instructions": "Does `answer` do what `rule` says?"}
}

A Choice picks one of the grades and returns a probability for each. A Noul is
the probability that a statement holds. The judgment comes back as numbers, so
the policy stays in code where I can see it and change it. One real answer from
the run:

verdict        correct 0.98 ยท partial 0.02 ยท wrong 0.00
unasked_extra  0.91  โ†’ at or above 0.9: "too much information" flag

About 2,000 tokens per answer, under a second, and all 188 answers for about three
cents.

Sonnet with the original PDF only sees the few claims Jev isn't sure about.

Keeping myself honest

I wrote the pass line down before I graded a second set of 20: at least 16 of
20, kappa at least 0.6.
That second set got used once, as a test, and never for
tuning.

It caught a mistake. On the first 20, I had tuned a penalty: if Jev was 90% sure
an answer added unasked details, the grade dropped to partial. It looked perfect
there. On the fresh set it caused two of the three disagreements. The judge
passed the line (17 of 20, kappa 0.63), but barely.

Reading those two answers showed why. I grade substance and length separately.
An answer can be right and too long. The detector itself was right: it flagged
exactly the four answers across both sets that I had found too long, and no
others. The penalty was the wrong policy. So it became a flag beside the grade.
Without it the judge agrees with me 19 of 20, kappa 0.85. That change came after
I'd seen the second set, so the next fresh set is where it has to hold.

The loop I'd run every time

The eval loop: golden set, all systems answer, I grade 20 blind, the judge grades the same 20 and I fix each disagreement, a fresh 20 against a line set in advance, then grade everything

I did it in a worse order: ran the systems, graded everything with an untested
judge, reported numbers, then calibrated. Next time:

  1. Build the golden set: questions as asked, reviewed gold answers.
  2. All systems answer, same documents, same instructions.
  3. I grade 20 answers blind.
  4. The judge grades the same 20. I read every disagreement and fix the cause, which is usually the answer key.
  5. Write the pass line down, then test on a fresh 20. Use it once.
  6. Only then grade everything and report.

A golden set isn't finished until a person has checked real answers against it.

What it means

On product datasheets, a frontier model with the PDFs attached is already very
good: my pipeline and the Claude app tied. Accuracy is the starting point. The
hard part is getting that accuracy into the place where buyers ask, across a
catalogue too big for one prompt, while datasheets change and the website drifts
away from them. Then proving it on the business's own questions, with a judge
calibrated to the person who answers them. That is what I'm building now.

The whole run cost about $8 in API calls. Most of that went on mistakes I won't
repeat. Now I estimate every paid run from a one-question trial, counting every
path that can call a paid model.

๐Ÿ“ฐ Read the original article on Dev.to AI

Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes โ€” full credit and traffic to the original publisher.