Dev.to AI 🤖 Ai 👁 0 📖 3 min read

I Built a Benchmark That Catches AI Models Cheating (And They All Failed)

I Built a Benchmark That Catches AI Models Cheating (And They All Failed) What task did you run? I built TwinBench, a benchmark that tests whether AI agents actually follow rules or just pattern-match their

I Built a Benchmark That Catches AI Models Cheating (And They All Failed)

What task did you run?

I built TwinBench, a benchmark that tests whether AI agents actually follow rules or just pattern-match their way to plausible-looking answers.

Here's the trick: every test item comes as a twin pair. Two stimuli that are identical except for one single controlling field. A model only gets credit if it answers both twins correctly. If it's just pattern-matching surface features, the twins will trip it up because the surface looks the same but the correct answer flips.

The benchmark covers three areas:

  1. Counterfactual Rulebooks (10 pairs): Config precedence, merge-queue ordering, timezone-aware batching, money rounding, journal replay, version ordering, authority resolution, retry gating. The kind of stuff that breaks production systems when agents get it wrong.

  2. Global Consistency Audits (10 pairs): Deployment provenance, event replay, snapshot lineage, change-freeze scoring, release shippability. Cross-referencing multiple sources of truth and finding contradictions.

  3. Temporal Reconciliation (6 pairs): Quota replays, FX cutovers, VOID cascades, DST payroll calculations, schema migration corrections. Time is hard, and agents need to handle it precisely.

Each pair uses a strict exact-match scorer. No partial credit. No "close enough." You either followed the rules correctly on both twins or you didn't.

Which models did you run it against?

Three frontier models, five runs each per pair:

  • Opus 5.5 Beta
  • Grok 4.6
  • GPT-6.1 Sol

I picked these three because they're the current top-tier reasoning models, and if any of them can handle precise rule-following under counterfactual pressure, it should be them.

The filtering protocol: a pair is retained (kept as a real benchmark item) only if at least one model fails at least 2 out of 5 runs. A run only passes if both twins get full credit. This ensures every item in the final benchmark is genuinely hard, not just hard for one model on a bad day.

What are the main insights?

Four pairs survived. All of them are brutal.

Pair Mechanism Opus Grok Sol
CR-N3 Timezone batching (DST vs fixed offset) 4/5 fail 4/5 fail 4/5 fail
CR-N5 Journal replay with ABORT semantics 5/5 fail 5/5 fail 5/5 fail
CR-N6 Version ordering (rc > release) 4/5 fail 5/5 fail 4/5 fail
CR-N8 Retry gating on idempotency registration 5/5 fail 5/5 fail 5/5 fail

CR-N5 and CR-N8: every model failed every single run. That's 30 consecutive failures across three different frontier models. These aren't edge cases, they're systematic blind spots.

What surprised me most:

  1. The models are confident and wrong. They don't hedge or express uncertainty. They produce well-formatted, detailed answers that are just... incorrect. The formatting is perfect, the reasoning sounds plausible, but they missed the actual rule.

  2. Twin pairs expose the cheating. On several items, a model would get twin A right and twin B wrong (or vice versa). If you're actually following the rule, the flip shouldn't matter. The fact that it does proves the model is pattern-matching, not reasoning.

  3. Shortcut baselines score zero. I built naive heuristic solvers (ignore timezone conversion, skip ABORT logic, always say "eligible"). They got 0/10 on every retained pair. This confirms the items can't be gamed with surface tricks, you actually need to do the work.

What I'd measure next:

  • More models: I'd love to test Gemini, Claude Sonnet, and open-weight models to see if the failure patterns are universal or architecture-specific.
  • Chain-of-thought analysis: Where exactly does the reasoning break down? Is it misreading the rule, misapplying it, or losing track across the twin?
  • The dropped pairs: Six CR pairs were too easy (models got 4-5/5 right). Are they easy because the mechanism is simple, or because the distractors weren't strong enough?

Where can we see it?

The benchmark harness, all 26 pair implementations, the filtering runner, shortcut baselines, and negative controls are all in the repo.

Built for the Kaggle Gemma 4 Developer Agent Paper Track. The twin-pair methodology is the core contribution: if your benchmark doesn't have counterfactual twins, you're measuring pattern matching, not reasoning.

📰 Read the original article on Dev.to AI

Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.