Dev.to AI 🤖 Ai 👁 0 📖 6 min read

My Eval Said RAG Made Things Up. My Eval Was Wrong.

I maintain django-explain-errors, a Django middleware that catches unhandled exceptions in development and asks an LLM to explain them. It works in two modes: RAG-off (default): the model sees only the Django tracebac

I maintain django-explain-errors, a Django middleware that catches unhandled exceptions in development and asks an LLM to explain them. It works in two modes:

  • RAG-off (default): the model sees only the Django traceback.
  • RAG-on: the model also sees relevant source code from your project, retrieved from a local sqlite-vec index by vector similarity.

The whole argument for RAG-on is grounding. If the model can read your actual code, it should stop guessing at function names and parameters. So I built an eval harness to test that claim, and the first results said the opposite: RAG-on appeared to fabricate more than RAG-off.

It took three versions of a single eval question to find out why. The short version: every layer in an eval pipeline can be working from less information than the thing it judges. When that happens, your metric measures the judge's blind spot, not the model.

The setup

The harness runs against a small, deliberately breakable Django blog app with 15 fixtures. Each fixture is a URL that triggers one realistic failure: a bad queryset lookup, a missing null check, a typo'd template name, a recursive __str__. Ten of them have their cause in the app's own code. Five have a cause the traceback already names.

For each fixture, the harness gets two explanations from gpt-4o-mini, one with RAG and one without. A separate judge model (Claude Sonnet via OpenRouter) compares them pairwise. It doesn't know which explanation used RAG, and which one it sees as "A" is randomized every time. The judge answers a fixed set of questions: did it find the cause, point to the fix location, propose a working fix, and explain it for a learner?

Notice what's missing from that list.

Version 0: nobody asked

I didn't ask about fabrication at first. But reading the judge's reasoning, I kept seeing it penalize explanations for "inventing" details, usually when the other questions came out tied. The judge was deciding fabrication on its own, as an unstated tiebreaker.

That meant fabrication was affecting the scores without showing up anywhere I could measure or audit. So I added it as an explicit question: no_fabrication.

Version 1: the judge had nothing to check against

The first no_fabrication judge saw the traceback and a short list of known facts about each fixture: the real cause and the real fix location. Its question was whether the explanation stated anything not supported by those.

RAG-on lost. It consistently got flagged for inventing details.

Then I looked at what was flagged. In the unexpected_kwarg fixture, RAG-on's explanation named a post_id parameter, and the judge called it fabricated. It wasn't. post_id was right there in the view's source, which RAG-on had retrieved and the judge had never seen.

The judge wasn't detecting fabrication. It was detecting specificity it couldn't verify. Any concrete detail that came from source looked invented, because the judge's only reference was the traceback. So the metric punished RAG-on for exactly the thing RAG is supposed to do well.

I reworded the criterion from "does not appear in" to "does not contradict," which helped at the edges. But the core problem remained: a judge that knows less than the generator will score the generator's extra knowledge as error.

Version 2: more context, same problem

The obvious fix was to give the judge the source. So I gave it the whole source modules involved in each failure.

It ignored them. The judge's answers barely changed, and its reasoning showed why: faced with 300 mostly irrelevant lines, marking a detail "unverified" is easier than finding the one line that settles it. The source was in the context window. It just wasn't being used.

Version 3: make it show its work

What finally worked was changing the shape of the task, not the amount of input.

The judge now gets a small, relevant source excerpt: the failing function itself, extracted by line number, plus any sibling methods it reaches directly. Instead of a yes/no on fabrication, it must list every checkable claim in each explanation and give each one a verdict:

  • verified: the excerpt supports it
  • contradicted: the excerpt says otherwise
  • absent: the excerpt doesn't cover it

no_fabrication is then derived from those verdicts in the parser, not asked as a direct question. The judge can't hand-wave anymore. Every "contradicted" points at a specific claim, which I can check.

With this in place, the result flipped. On the 10 fixtures whose cause lives in app code, RAG-on avoided fabrication in 26 comparisons against RAG-off's 17. On the other 5, it was 14 against 5. RAG-on hadn't been fabricating more. The judge had been unable to check its work.

That second number surprised me in a different way. I designed those 5 fixtures as a control: the traceback already names the cause, so RAG shouldn't matter. It mattered more there than anywhere. The eval was testing my assumptions, not just the model.

What RAG-off fabricates is revealing. Without source, gpt-4o-mini tends to invent a plausible function signature rather than say it doesn't know. In one run, RAG-off confidently wrote def post_preview(request): when the real signature is def post_preview(request, post_id). In another, it said a template needed {% load humanize %} added when line 1 already had it.

The bug hiding in the numbers

While chasing fabrication, I found a second problem affecting a different metric.

RAG-on was winning points_to_fix_location by a wide margin. But the package truncated long tracebacks by keeping only the tail. For a Django ORM error, the tail is mostly library internals, so the frame naming the failing app function routinely got cut. RAG-off often never saw where the error happened.

I didn't find this by looking at scores. I found it by looking at what RAG-off was actually sent.

After fixing truncation to preserve app frames, RAG-off's fix-location score on app-code fixtures rose from 6 of 30 to 13. RAG-on stayed roughly where it was. About half of the original gap was a truncation bug. The other half is real. And because the truncation lives in the middleware, the fix didn't just correct the eval. It made the package better for everyone using it, with or without RAG.

What I still don't trust

The final numbers across three runs: RAG-on won 35 of 43 scored comparisons, RAG-off 5, with 3 ties. A full three-run pass costs about $1.37, almost all of it the judge writing out claim lists. But there are limits I can't wave away:

  • Overlap bias. The judge checks claims against the failing function's source, which is the same source RAG-on retrieves from. Part of RAG-on's fabrication advantage may come from the judge and generator sharing material RAG-off never sees, not from RAG-on being more careful.
  • Spot-checks are a sample. A script flags claims marked "contradicted" that still mention real identifiers from the source. On one run it flagged 14 of 363 claims, and all 14 verdicts held up on manual review. That's a sample with no false positives, not proof there are none.
  • Retrieval can mislead. In the missing_post_key fixture, RAG-on anchored on an adjacent retrieved template and pointed the fix there instead of at the view. Retrieval reduces fabrication. It doesn't eliminate wrong answers.
  • Stale numbers. I later fixed a source-extraction bug affecting one fixture, and I haven't re-run the full suite since.

Takeaways

  1. Ask every judged question explicitly. If a criterion matters, the judge is already scoring it. Make it visible so you can measure and audit it.
  2. Make the judge show its work. Claim-level verdicts, with the score computed in code, can be checked line by line. A bare yes/no can't.
  3. Inspect inputs and outputs, not just scores. Every fix in this story came from reading something other than a number: the judge's reasoning, the flagged claims, the prompt RAG-off actually received. Check what each layer can see, because a gap between what the generator saw and what the judge saw looks exactly like model behavior.
  4. When a result surprises you, suspect the instrument first. "RAG fabricates more" wasn't a finding about RAG. It was a finding about my judge.
  5. The most valuable output might not be the score. This eval produced a headline number, but its biggest payoff was what it exposed along the way: a broken metric, a truncation bug that was degrading explanations for every user, and a control group that wasn't one. A score tells you where you are. Building the instrument tells you what you got wrong.

The harness, fixtures, judge prompt, and full results are in the repo under evals/, with the history of these changes in evals/METHODOLOGY.md.

📰 Read the original article on Dev.to AI

Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.