Dev.to AI πŸ€– Ai πŸ‘ 0 πŸ“– 10 min read

Invented O'Clock: I removed one fact from 45 scheduling problems to see which AI models make up a time

This is a submission for the Kaggle Benchmarking Challenge What I Benchmarked Ask an assistant a scheduling question and you usually get a crisp answer: leave at 7:40 AM, the call is at 4:30 PM their time, th

This is a submission for the Kaggle Benchmarking Challenge

What I Benchmarked

Ask an assistant a scheduling question and you usually get a crisp answer: leave at 7:40 AM, the call is at 4:30 PM their time, the invoice is due November 14. The trouble is that real questions often leave out one fact the answer depends on: when the train departs, which time zone the other person is in, the date the 30-day clock started. Without that fact there is no correct time to give. A careful person says "I can't tell without X." A model trying to be helpful often fills the gap with a plausible guess, in exactly the same confident format as a computed answer.

In a chat window that's an annoyance. In scheduling and ops tools it's a bug. An invented time can end up in a calendar invite, a reminder, a shift roster or a payment due date, and nothing downstream can tell it apart from a real one. unknown is a useful answer because a tool can act on it: ask the user, hold the action, flag it for review. So it's worth measuring two things, not one: whether a model can do the time math, and whether it notices when the math can't be done.

So I built Invented O'Clock, a benchmark that asks two questions about the same everyday scheduling problems:

  1. Can the model do the arithmetic? Leave-by times, cooking backwards from dinner, overnight shifts, time zones in the weird weeks when the US and Europe are out of sync on daylight saving, net-30 due dates, business days around holidays, "the first Tuesday after the first Monday", countdowns and recurring events.
  2. Does it notice when it can't? Every problem has an evil twin with one key fact removed. The right answer there is unknown, and anything else is a time or date the model made up.

Here's one pair. The answerable twin:

On Friday, March 19, 2027, Maddie has a call at 10:30 AM Chicago time. The client she is calling is in Berlin. The call is booked for 45 minutes.
Question: What is the local time for the client when the call starts?

(4:30 PM. The US has already switched to daylight time and Europe hasn't, so the usual 7-hour gap is 6 for two weeks.)

The missing-fact twin is the same text without the bolded sentence. The only correct final answer is unknown.

Design choices:

  • 45 scenarios Γ— 2 twins = 90 items across 9 families (5 each). Every prompt offers ANSWER: unknown as a legal option, word for word, so nobody is tricked into answering.
  • Two kinds of missing fact. In 27 twins the gap is flagged ("Dana doesn't know how long the drive takes"). In 18 it is silent: the fact is simply gone. This separates "can't read" from "doesn't notice".
  • Distractors on purpose. Each prompt includes a time or date that is not the answer (an alarm time, a kickoff date). The grader can tell three wrong behaviors apart: inventing a new value, echoing a value from the prompt, and refusing a question that was answerable.
  • Deterministic grading, no LLM judge. The grader reads the final ANSWER: line and parses times and dates with regex and Python's datetime. A date answer with the wrong weekday counts as wrong.
  • Trustworthy answer key. Answers are computed by two independent implementations (one uses datetime/zoneinfo, the other uses Julian Day Numbers and hand-entered UTC offsets) and spot-checked by hand. A 71-test suite checks that the twins differ only in the key fact, that no missing twin leaks the fact back in, and that the grader separates 7 simulated "models" (perfect, always-unknown, off-by-a-bit, echo-a-distractor, invent-on-missing, and so on).

On Kaggle it's two tasks grouped into one benchmark: invented_oclock_solve (accuracy on the 45 answerable twins) and invented_oclock_abstain (share of the 45 missing-fact twins answered unknown). Neither score means much alone. A model that always says "unknown" gets 100% on abstain and 0% on solve, so you have to read them together.

Models Tested

I ran nine models on both tasks: Claude Haiku 4.5, Claude Sonnet 4.6, Gemini 3.1 Flash-Lite (preview), Gemini 3.1 Pro (preview), Gemini 3.7 Flash, Gemma 4 26B A4B, Qwen3 235B A22B Instruct, GLM-5, and DeepSeek-R1 0528. The mix covers Anthropic (small and large), Google (lite, flash, pro, and the open Gemma), plus Alibaba's Qwen, Zhipu's GLM-5, and one explicit reasoning model (DeepSeek-R1). Every model got the same setup: one answer per item, each provider's default temperature (the Kaggle SDK doesn't send a temperature to its model proxy), and a cap of 4,096 output tokens per answer, hidden reasoning included.

A note on errors (methods). My first round of runs (v1) was mostly broken, and the cause was my setup, not the models. I launched eight models at once against Kaggle's free $10/day quota. Kaggle reserves quota for a model's maximum possible output before each call (up to about $0.96 per call), so most calls were refused with a "quota exceeded" error before they ever reached a model, and a few runs hit a timeout that killed the whole run. I discarded that round, and none of the numbers in this post come from it. For the v2 runs reported here, I capped output at 4,096 tokens, ran two items at a time (n_jobs=2), left temperature at each provider's default, and retried any item that hit an API error once. An item that still errored was not graded: it's left out of that model's denominator and counted in a separate "Errored" column, never scored as a wrong answer. A run with more than 2 errored items reports no score and is re-run. In the final v2 set, all nine models finished 45/45 on both tasks with 0 errored items.

Two things in the logs are worth knowing before you read the tables:

  • Token cap. A response that ran into the 4,096-token cap was graded normally: if it never reached a final ANSWER: line, it fails. DeepSeek-R1 hit the cap on 24 responses (14 solve, 10 abstain) and GLM-5 on 6. More on what that means in the findings.
  • Empty proxy messages. The Gemma 4 26B logs, and GLM-5's too, include warnings that Kaggle's model proxy returned an empty message (choices[0].message=None), which the SDK treats as an empty response. Both runs still completed 45/45 with 0 errored items, and every graded item in their result CSVs has a non-zero output-token count, but I can't map the warnings to specific items from the logs alone.

Findings

Overall

Model Solve (answerable) Abstain (missing fact) Abstain when gap is flagged Abstain when gap is silent Invented a value Used a value from the prompt Both twins right Errored, excluded (solve / abstain) Hit the token cap
anthropic/claude-haiku-4-5@20251001 82% (37/45) 84% (38/45) 93% (25/27) 72% (13/18) 7 0 71% (32/45) 0 / 0 0
anthropic/claude-sonnet-4-6@default 91% (41/45) 89% (40/45) 96% (26/27) 78% (14/18) 5 0 84% (38/45) 0 / 0 0
deepseek-ai/deepseek-r1-0528 69% (31/45) 80% (36/45) 81% (22/27) 78% (14/18) 9 0 58% (26/45) 0 / 0 24
google/gemini-3.1-flash-lite-preview 93% (42/45) 82% (37/45) 85% (23/27) 78% (14/18) 8 0 76% (34/45) 0 / 0 0
google/gemini-3.1-pro-preview 100% (45/45) 96% (43/45) 93% (25/27) 100% (18/18) 2 0 96% (43/45) 0 / 0 0
google/gemini-3.7-flash 100% (45/45) 98% (44/45) 96% (26/27) 100% (18/18) 1 0 98% (44/45) 0 / 0 0
google/gemma-4-26b-a4b 82% (37/45) 91% (41/45) 93% (25/27) 89% (16/18) 2 0 76% (34/45) 0 / 0 0
qwen/qwen3-235b-a22b-instruct-2507 87% (39/45) 82% (37/45) 89% (24/27) 72% (13/18) 8 0 73% (33/45) 0 / 0 0
zai/glm-5 91% (41/45) 96% (43/45) 93% (25/27) 100% (18/18) 0 0 87% (39/45) 0 / 0 6

Percentages are over graded items only. Items that hit an infrastructure or quota error (after one retry) were not graded: they are excluded from every denominator and counted in the "Errored" column. "Both twins right" = solved the answerable twin AND said unknown on its missing-fact twin, over scenarios where both twins were graded. "Hit the token cap" = graded responses that used all 4096 output tokens (graded normally).

By scenario family (solve / abstain, graded items, out of 5 each when nothing errored)

Model Leave-by Cook backwards Finish time Time zones Invoice terms Business days Nth weekday Countdown Recurring
anthropic/claude-haiku-4-5@20251001 5/5 / 5/5 4/5 / 3/5 4/5 / 5/5 2/5 / 5/5 4/5 / 4/5 5/5 / 5/5 5/5 / 4/5 3/5 / 5/5 5/5 / 2/5
anthropic/claude-sonnet-4-6@default 5/5 / 5/5 4/5 / 4/5 5/5 / 5/5 5/5 / 5/5 5/5 / 4/5 5/5 / 5/5 3/5 / 5/5 5/5 / 5/5 4/5 / 2/5
deepseek-ai/deepseek-r1-0528 5/5 / 5/5 5/5 / 4/5 5/5 / 4/5 5/5 / 4/5 1/5 / 5/5 5/5 / 5/5 4/5 / 3/5 1/5 / 4/5 0/5 / 2/5
google/gemini-3.1-flash-lite-preview 5/5 / 5/5 4/5 / 5/5 5/5 / 5/5 4/5 / 5/5 5/5 / 4/5 5/5 / 5/5 5/5 / 3/5 4/5 / 4/5 5/5 / 1/5
google/gemini-3.1-pro-preview 5/5 / 5/5 5/5 / 5/5 5/5 / 5/5 5/5 / 5/5 5/5 / 5/5 5/5 / 5/5 5/5 / 4/5 5/5 / 5/5 5/5 / 4/5
google/gemini-3.7-flash 5/5 / 5/5 5/5 / 5/5 5/5 / 5/5 5/5 / 5/5 5/5 / 5/5 5/5 / 5/5 5/5 / 4/5 5/5 / 5/5 5/5 / 5/5
google/gemma-4-26b-a4b 5/5 / 5/5 5/5 / 5/5 5/5 / 5/5 5/5 / 5/5 4/5 / 5/5 4/5 / 5/5 3/5 / 4/5 2/5 / 5/5 4/5 / 2/5
qwen/qwen3-235b-a22b-instruct-2507 4/5 / 5/5 5/5 / 3/5 5/5 / 5/5 4/5 / 5/5 5/5 / 4/5 5/5 / 5/5 5/5 / 4/5 3/5 / 5/5 3/5 / 1/5
zai/glm-5 5/5 / 5/5 5/5 / 5/5 5/5 / 5/5 5/5 / 5/5 4/5 / 5/5 5/5 / 5/5 4/5 / 4/5 3/5 / 5/5 5/5 / 4/5

Finding 1: Two Gemini models nearly nail both twin checks. Gemini 3.7 Flash scored 100% on solve (45/45) and 98% on abstain (44/45), with both twins right on 44 of 45 scenarios and only 1 invented value. Gemini 3.1 Pro is right behind: 100% solve (45/45), 96% abstain (43/45), both twins right on 43 of 45, 2 invented values. Next is GLM-5 at 87% both-right (39/45) with 0 invented values. That's the bar in this set: arithmetic and restraint together, not one without the other.

Finding 2: Silent gaps are harder than flagged ones, for most models. For six of the nine models, abstain drops when the missing fact is simply absent instead of announced: Haiku falls from 93% flagged to 72% silent, Sonnet from 96% to 78%, Qwen from 89% to 72%, Flash-Lite from 85% to 78%, Gemma from 93% to 89%, and DeepSeek-R1 from 81% to 78%. The other three (Gemini 3.1 Pro, Gemini 3.7 Flash, GLM-5) score 100% (18/18) on silent gaps.

Finding 3: When models fail abstain, they invent. They never echo. Across all nine models, 42 of the 405 missing-fact responses were graded as inventing a time or date (DeepSeek-R1 9, Flash-Lite 8, Qwen 8, Haiku 7, Sonnet 5, Gemini 3.1 Pro 2, Gemma 2, Gemini 3.7 Flash 1, GLM-5 0). "Used a value from the prompt" is 0 for every model. The planted distractor times almost never got copied; the wrong answers are new values.

Finding 4: DeepSeek-R1's score is partly a token-budget result. DeepSeek-R1 0528 finished last: 69% solve (31/45), 80% abstain (36/45), both twins right on 58% (26/45). But it hit the 4,096-token cap on 24 of its 90 responses (14 solve, 10 abstain), and every one of its 23 failed items was a capped response. On solve, all 14 misses ran to the cap: 10 never reached an ANSWER: line, 3 ended in unknown on an answerable question, and 1 was unparseable. On abstain, all 9 "invented" responses also ran to the cap. Under my grader, a response with no final ANSWER: line that mentions a time or date not in the prompt counts as invented, so some of those 9 may be reasoning cut off mid-calculation rather than a committed guess. One capped abstain response still passed. These were graded normally, as the methods say, so read this row as "DeepSeek-R1 with 4,096 output tokens, reasoning included", not as its ceiling. The same thing shows up on a smaller scale elsewhere. All 6 GLM-5 misses (4 solve, 2 abstain) were capped responses. Gemma 4 26B shows 0 in the cap column, but all 10 of its "no answer line" failures (8 solve, 2 abstain) stopped at 4,092 to 4,093 tokens, just under the 4,096 the table counts, so they look budget-limited too.

Finding 5: Recurring and countdown are the weak families. Recurring abstain is 1/5 for Flash-Lite and Qwen, and 2/5 for Haiku, Sonnet, Gemma and DeepSeek-R1. Countdown solve is 1/5 for DeepSeek-R1, 2/5 for Gemma, and 3/5 for Haiku, Qwen and GLM-5. DeepSeek-R1 also went 0/5 on recurring solve and 1/5 on invoice-terms solve, all capped responses. Leave-by and business days are near the ceiling for everyone: on missing-fact twins every model went 5/5 in both families, and on answerable twins every model went 5/5 except Qwen (4/5 leave-by) and Gemma (4/5 business days).

What surprised me: Solve skill and restraint don't track each other. Gemini 3.1 Flash-Lite solves 93% (42/45) but invented 8 values, while Gemma 4 26B solves only 82% (37/45) and invented 2. The other surprise was that the one explicit reasoning model came last, and the logs say a big part of that is running out of room, not getting the math wrong.

What it changed about how I read these models: A solve score on its own says nothing about whether a model will make up an answer when it can't know one, so any tool that writes times into calendars or reminders should test the missing-fact case too. And for reasoning models, the output budget is part of the result. A cap that's fine for one model can quietly turn another model's long reasoning into wrong answers.

Limitations. 45 scenarios is small: one item is about 2 points, so treat gaps under ~10 points as noise. Each model ran once, at its provider's default temperature, so a re-run can shift a few items. The scenarios are US-centric, in English, and written by one person. The strict format check means a model that hedges on its final line ("unknown, but probably 9:05 AM") counts as inventing. That's deliberate, but it is a choice. The 4,096-token cap penalizes long-reasoning models, as Finding 4 shows.

What I'd measure next: DeepSeek-R1 and the other capped models again with a much larger output budget, to separate "ran out of room" from "got it wrong"; the same twins in a multi-turn chat where the user pushes back ("just give me your best guess"); letting models use a Python tool; and a version where the missing fact appears earlier in the conversation instead of being absent.

My Benchmark

Everything is self-contained in the task notebooks (items, grader, scoring), so you can fork either task and run it on any model Kaggle offers.

AI-use note: The benchmark code, the 90 items, and this post were built with AI assistance (Grok Bot) and reviewed by me. Every answer key is computed by two independent solvers and spot-checked by hand, and every number in this post comes from the v2 Kaggle run logs for the tasks linked above.

πŸ“° Read the original article on Dev.to AI

Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes β€” full credit and traffic to the original publisher.