It scored 100%. Its note scored 0%
This is a submission for the Kaggle Benchmarking Challenge What I Benchmarked I kept wiping the chat. A model sees a hidden rule through examples. Four rounds, six items each, with the right answers after ev
This is a submission for the Kaggle Benchmarking Challenge
What I Benchmarked
I kept wiping the chat.
A model sees a hidden rule through examples. Four rounds, six items each, with the right answers after every round. Then it has 120 words. Then the conversation is deleted. A fresh reader gets the note and a set of questions it has never seen.
Each dot below is one of those episodes. Left to right: how the writer scored while the examples were still on screen. Bottom to top: how the fresh reader scored from the note alone. The orange dots are the ones that bothered me. The writer was at least 90 percent right. The reader, holding only the note, landed at 50 percent or worse.
The rules are made up. Strings, numbers, a tiny grammar, categories. Three difficulties each, one episode per cell, same grid for every model. The generator holds the answer key, so a model never grades another model.
Under every note I also ran two controls on the same questions. No note at all. And a perfect note, written from the true rule, capped at the same 120 words. Teaching Efficiency is the share of that gap the model's own note managed to close. A 1.0 means the note did what the truth did. A 0 means the reader would have done as well with a blank page. Under 0, the note pushed the reader the wrong way. A run that died halfway is left blank. It is not a zero.
For that teaching score the fresh reader is always Gemini 3.1 Flash-Lite. I wanted the note to work for a small model. A note that only its author can use is just a diary.
After that I flipped the setup. Every model read the whole pile, 320 notes, plus the perfect note and the blank, and answered the same questions. That score is Reading Efficiency. I only compare readers when the perfect note actually works for them.
Models Tested
Twenty-five models finished the teaching run. I wanted three kinds of company in the same table.
Neighbours. Sonnet 4.5 beside Sonnet 5. Opus 4.8 beside Opus 5. Gemini from 2.5 Flash up through the 3.x Flash line to Gemini 3.1 Pro.
An open model in with the hosted ones. Gemma 4 31B sits on the same chart as GPT-6 Astra and Gemini 3.5 Flash.
And the small ones that usually get left out of a write-up. GPT-5.4 nano, GPT-5.4 mini, both Flash-Lites, Gemma 4 26B.
Two names are missing, and I am not filling them in. gpt-oss-120b never got through a run. The provider was overloaded. Grok 4.20 non-reasoning rejected the reasoning setting everyone else used. They stay off the chart.
Findings
This is the teaching board. The bar is how much of a perfect note the model's own note was worth. Gemini 3.5 Flash is at the top, 1.02. Sonnet 5 is one of the short ones, 0.75, under several models that cost less and learned less.
Gemma 4 31B reached 99 for $0.15. Gemini 3.5 Flash is the only model that spent more and still lifted the score, to 102 at $0.78. GPT-6 Astra spent about a dollar to land where Gemma already was.
The note failed its own examples
Sonnet 5 learned the set at 0.92 and taught it at 0.75. On one number rule the truth is 4x β 5. The last practice rounds were perfect. Then it wrote that the rule was 4x β 1, listed 57β223, 1ββ1, and 25β95, and said the examples match. 4Γ57 β 1 is 227, not 223. The fixed reader scored 0 out of 12. So did a fresh Sonnet 5, handed only that note. The rule was in the chat. It did not survive the paragraph.
Opus 5 went the other way. It learned at 0.78, the weak end of the frontier, and its notes scored 0.99. On both easy string rules the last practice round was 0 for 6. The note still got the small reader to 12 for 12. One note said to reverse the word. The other said where to move the first letter. A fresh copy of Opus, reading its own note, came out at 1.04 against the perfect note.
Ten episodes in the full set look like Sonnet's. The writer finishes practice at 90 percent or better. The reader, with the note, is at 50 percent or worse. Nano learned to reverse a word at 100 percent and described an interleave. The reader scored 0, below the blank-page score. A mini grammar note announced the wrong word order, then used the right order in its own example. The reader scored 0. DeepSeek-R1 learned a string rule cold. The small reader got 46 percent from the note. DeepSeek's own fresh copy got 100 percent from the same page.
Some notes are still thinking
Ten notes still have the scratch work in them. "Wait." "Actually." "Let's look." Those ten teach at 0.47. The other 288 teach at 0.91. On the four where the writer had already learned the rule, the scratch-work notes teach at 0.33. The notes that just say the rule teach at 0.97.
Flash-Lite had a reversal at 92 percent, then talked itself into a vowel swap. The reader scored 8 percent. With a blank page that same reader scored 31 percent.
Extra examples did little. Notes with a pair of worked examples taught at 0.90. Notes without them taught at 0.88. Shorter was fine. Gemma 4 31B averaged about 44 words and taught at 0.99. Sonnet 4.5 averaged about 94 and taught at 0.92. Sonnet 5 averaged about 65 and taught at 0.75.
Three notes landed more than ten points under the blank page. That is rare. The usual miss is a sure sentence parked next to examples that disagree with it.
A good note, and the readers bunch up
Each row is a writer. Each column is a reader who never saw the lesson, only the note. The last two rows are the perfect note and the blank page. Opus 5's column is the washed-out one. Those are empty replies.
Where the note is solid, the readers land in a tight band. Opus 4.8, Grok 4.20, and GPT-5.5 each got about 93 percent of the model-written notes. Gemini 3.6 Flash and Gemini 3.1 Pro got about 92 and 93. The perfect note was essentially 100 percent for them. Gemini 3.5 Flash, best teacher in the whole run at 1.02, read other people's notes at about 92 percent. Reading Efficiency 0.90. Ordinary.
Gemma 4 31B taught at 0.99 and the run cost $0.15. Astra taught at 1.01 and the run cost $0.88. Gemma's notes were the short ones. As a reader it scored 91 percent on everyone else's notes.
Sonnet 5, author of the 4x β 1 note, is better at other people's notes than at its own. Its notes, read by itself, scored 0.75. Opus 5's notes, read by Sonnet 5, scored 0.99.
The ratio will crown the wrong reader
Green is the perfect note. Blue is whatever the model actually wrote. Orange is a blank page. The blue line is only a fair grade where the green line is still up near 1.
Gemma 4 26B's Reading Efficiency is 0.96. That looks like the top of the table. On the notes themselves it scored 80 percent. The perfect note only got it to 83 percent. With nothing, it scored 0. The ratio is large because the bottom of the fraction is empty. Qwen3 Next 80B rhymes with it: ratio 0.95, notes at 80 percent, perfect note at 83 percent.
Opus 5, reading, is a different kind of wreck. Empty replies on 9 of 12 perfect notes, and on 129 of 296 model notes. Fifty percent on the notes. Twenty-two percent on the perfect note. A modest numerator over a broken ceiling comes out above 2. I am not putting that on a leaderboard. An earlier probe hit the same silence when one sentence about a hidden rule was in the prompt.
Nano's 0.93 is a smaller version of the Gemma problem. It scores about 6 percent with no note, so the ratio looks sharp. On the notes it is at 88 percent, middle of the pack.
Flash-Lite is what a low score looks like when the perfect note still works. Eighty-five percent on the notes, 16 percent with nothing, 99 percent on the perfect note, Reading Efficiency 0.83. That 0.83 can be compared with the others. Gemma's 0.96 needs the three raw rates sitting next to it, or it will be misread.
What I would keep
Print the blank-page score and the perfect-note score beside the ratio. A teaching score of 1.0 means the note matched the truth. A reading score of 0.96 can mean the reader was lost with nothing and only middling with a perfect note.
Grade the note with a stranger. DeepSeek's own fresh copy scored 100 percent on a note that the small reader scored 46 percent on, after DeepSeek had learned the rule completely.
Count the blank page. Three notes did worse than no note. They added a rule that was false.
Leave empty replies and dead runs out of the average. Opus 5's blanks inflated a ratio. gpt-oss never finished. Those two facts do not belong in the same cell.
What I would measure next
Hand the note to a second model and make that model write the next note. Same rules. Count how many hops the rule lasts. Sonnet's "the examples match," written over examples that do not match, is the sentence I want a model to catch before it passes the page on.
My Benchmark
Teaching and reading are separate tasks. Teaching Efficiency is the first number. Reading Efficiency is the second, and a reader whose perfect note failed is marked not comparable. Twenty-five models have a teaching score. Twenty-four finished reading. Sonnet 4.5 has a teaching score. Its reading run was still going when I wrote this. gpt-oss-120b and Grok 4.20 non-reasoning are not in the scored rows.
Teaching: https://www.kaggle.com/benchmarks/tasks/naomiandreapereira/amnesiac-teach
Reading: https://www.kaggle.com/benchmarks/tasks/naomiandreapereira/amnesiac-read
Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes β full credit and traffic to the original publisher.




