Dev.to WebDev 🛠 Dev 👁 0 📖 5 min read

Don't let the LLM do the maths: grading writing with AI but scoring in code

If you ask a language model to grade a text and also to count the words, add up the points and apply the scoring rules, it will usually give you a confident answer. Sometimes the answer will be wrong, and you won't notic

If you ask a language model to grade a text and also to count the words, add up the points and apply the scoring rules, it will usually give you a confident answer. Sometimes the answer will be wrong, and you won't notice, because the output looks reasonable.

I ran into this while building the writing and speaking practice in ExamReady, a platform for preparing for German B1/B2 exams. The setup I ended up with is simple to state: the LLM judges, and my code does everything that can be computed. This post explains where that line is and why I drew it there.

ExamReady is an independent project and is not endorsed by ÖSD, ÖIF or the Goethe-Institut.

The problem

In a writing exam, a learner gets a task, writes a text, and is graded on several official criteria. There are also hard rules around the result. One example: if the text has less than 50% of the required number of words, it scores 0 points.

That rule has nothing to do with language quality. It's arithmetic. A text can be well written and still score 0 because it's too short.

The tempting approach is to put the whole grading guide into one prompt and let the model produce a final score. That's one API call and one result, which is attractive when you're a solo developer. But it mixes two very different jobs.

What models are good at, and what they aren't

A model is good at judging things that need language understanding. Does this text address the task? Is the register appropriate? Is the grammar mostly correct, and where does it go wrong? Those are the questions I want an LLM for.

A model is unreliable at counting and at arithmetic. Counting the words of a text, deciding whether that is above or below half the required number, and summing criterion scores into a total are all things that a few lines of code do exactly, every time. Asking a model to do them gives you a result that is usually right, and occasionally wrong in a way that is hard to detect and impossible to reproduce.

So I split the work.

The split

The LLM (Gemini) gets the task and the learner's text, and returns grades per official criterion, plus written feedback. It does not return a total and it does not decide whether the text was long enough.

The code then:

  1. counts the words in the learner's text
  2. applies the official rules, for example 0 points if the text is under 50% of the required words
  3. combines the criterion grades into the final result

Here is the shape of the model output I ask for:

type CriterionGrade = {
  criterion: string;     // one of the official criteria
  points: number;        // within that criterion's allowed range
  comment: string;
};

type ModelGrading = {
  grades: CriterionGrade[];
  feedback: string;
  corrections: { original: string; suggestion: string }[];
};

And the part that belongs to the code:

function countWords(text: string): number {
  return text.trim().split(/\s+/).filter(Boolean).length;
}

function finalScore(text: string, task: WritingTask, g: ModelGrading) {
  const words = countWords(text);

  // rule applied in code, not by the model
  if (words < task.requiredWords * 0.5) {
    return { total: 0, reason: "too_short", words };
  }

  const total = g.grades.reduce((sum, c) => sum + clamp(c.points, task), 0);
  return { total, words };
}

Two small details matter here. First, the word count is a function I can unit-test. I can feed it texts and check the result, which I can't do with a prompt. Second, the model's points are clamped to the allowed range for each criterion, because a model can return a value outside what the scale permits. I treat model output as input that needs validation, not as a verdict.

Verifying corrections against the learner's text

The second place where I stopped trusting the model blindly is corrections. A useful feedback feature is showing the learner a specific mistake in their own sentence and a better version. The failure mode is that the model quotes something the learner never wrote, or slightly rewrites their sentence before "correcting" it.

So every correction is verified against the learner's own text before it is shown. In essence:

function verifiedCorrections(text: string, list: ModelGrading["corrections"]) {
  return list.filter((c) => text.includes(c.original));
}

If the quoted passage isn't actually in the text, the correction is dropped. A missing correction is better than a wrong one pointing at a sentence the learner didn't write. The real check can be a bit more tolerant about whitespace, but the principle is the same: the code checks the model's claims against the source.

Why not do everything in the prompt?

I can see the argument for it. Modern models are good, and often they'll count correctly. But "often" isn't the standard I want for something that decides whether a learner thinks they would pass. A few reasons I prefer the split:

  • Reproducibility. The same text always gets the same word count and the same rule application. Only the judgement part varies.
  • Testability. Rules in code can have tests. Rules in prompts have vibes.
  • Debuggability. When a score looks wrong, I can tell whether the rule or the model's judgement caused it.
  • Rule changes. If a rule changes, I change a function, not a prompt I then have to re-validate.

The cost is more code and a more structured output from the model. I think that's cheap for what it buys.

A note on speaking

Speaking practice uses a live AI conversation partner built on Gemini Live. It's a different shape of problem, since it's a conversation rather than a document, and it costs me roughly ten cents per speaking task. I'm still learning what the right boundary is there, so I'll keep this post about writing.

Takeaways

  1. Let the LLM judge language, and let code count, compare and add up.
  2. Treat model output as untrusted input: validate ranges and shapes.
  3. Apply official rules in code so they can be tested and reproduced.
  4. Verify any quote or correction against the learner's own text before showing it.
  5. Prefer a missing piece of feedback over a wrong one.

None of this is clever. It's mostly about not asking a language model to do what a function does better.

I'm building this at examready.at, feedback welcome.

📰 Read the original article on Dev.to WebDev

Originally published by Dev.to WebDev. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.