Can LLMs Keep Score? I Benchmarked Tennis and Padel Score-Tracking
This is a submission for the Kaggle Benchmarking Challenge What I Benchmarked I wanted a test where the right answer is unambiguous, so I picked something simple that is surprisingly easy to get wrong: keepin
This is a submission for the Kaggle Benchmarking Challenge
What I Benchmarked
I wanted a test where the right answer is unambiguous, so I picked something simple that is surprisingly easy to get wrong: keeping score in a match.
Each case gives a model a list of points in order ("A A B A B ...", each letter is who won that point). The model has to report the state right after the last point: sets, games, points in the current game, and who serves next. I wrote a small scoring engine in Python that computes the true answer, so every reply is simply right or wrong. No judge model is involved.
The benchmark covers tennis (advantage scoring) and padel (golden point). It includes sequences of 30, 60 and 100 points, plus special cases forced into a 6-6 tiebreak, where the serving order is awkward. Every point sequence is asked twice: once with the full rules written in the prompt, and once with only "Sport: padel (golden point scoring...)", so the model has to rely on what it already knows. The final benchmark has 48 cases.
Models Tested
I ran it on five models available in Kaggle Benchmarks: Gemini 3.7 Flash, gpt-oss-20b, Claude Haiku 4.5, GPT-5.4 nano, and Gemma 4 26B A4B (this one errored on Kaggle, so it has no score). I chose a mix of a strong fast model, a small reasoning model, and small models that answer quickly, because I expected the score to depend on how much a model works things out step by step. Everything ran with default settings, once per case.
Findings
| Model | Correct Cases | Accuracy | Status |
|---|---|---|---|
| Gemini 3.7 Flash | 46/48 | 96% | Completed |
| GPT-OSS-20B | 20/48 | 42% | Completed |
| Claude Haiku 4.5 | 0/48 | 0% | Needs investigation |
| GPT-5.4 nano | 0/48 | 0% | Needs investigation |
| Gemma 4 26B A4B | — | — | Evaluation error |
- The spread is huge. Scoring a game of tennis is easy for a person, and two models got none of 48 right.
- Not working it out vs working it out badly. GPT-5.4 nano wrote about 25 tokens per case, just the answer line, and guessed. Some of its answers were impossible states, like a match at 2-1 in sets when the match was still going. Claude Haiku 4.5 wrote about 1,100 tokens per case, counting points in plain text, and still scored 0.00. In the one reply I read closely, it contradicted itself about who had won which games and then reported a state that didn't match its own working. So writing out the counting isn't enough on its own.
- A lot of the errors I saw were noise. In an earlier 100-case development run, Gemini 3.7 Flash scored about 95%, and all 5 of its mistakes were in the no-rules condition. That looked like "the model needs the rules." I reran those 5 cases twice, and 9 of 10 passed. The same model on the same prompt gave different answers, so single-run results overstate differences between conditions.
- I did not find a padel weakness for Gemini. Padel and tennis scores were within one error of each other, which is not a real difference at this sample size.
What I'd Measure Next
- Run each case several times, because one run per case is too noisy to compare conditions.
- Turn on extended reasoning for the models that scored 0.00, to see whether the problem is the model or the default setting.
- Get per-sport and per-length results for every model, not just Gemini.
- Try 200-point sequences, since the best model was close to the ceiling.
Limits
48 cases is small. I only read a few of Claude Haiku's replies, so I can't rule out that some long ones were cut off. The padel rules (golden point, serving order after a tiebreak) come from my reading of the rules and I haven't checked them against the official rulebook.
My Benchmark
https://www.kaggle.com/benchmarks/gulrezqayyum/tennis-and-padel-state-tracking-benchmark
Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.
