Dev.to AI 🤖 Ai 👁 0 📖 5 min read

Is a 10x Jump Always a 10x Jump? A Statistical Audit of the METR Time-Horizon Plot

If you follow AI progress, you have seen the METR plot: a line on a log axis showing that the length of software tasks AI systems can complete keeps growing exponentially. A new paper from two Berkeley statisticians, On

If you follow AI progress, you have seen the METR plot: a line on a log axis showing that the length of software tasks AI systems can complete keeps growing exponentially. A new paper from two Berkeley statisticians, On the estimation and validity of AI time horizons (Nguyen and Fithian, October 2026), does not dispute the trend. It asks a narrower question that matters to anyone who uses benchmarks to make decisions: what exactly is the y-axis, and is a "10x longer horizon" the same kind of improvement everywhere on the axis?

What the time horizon actually measures

The metric comes from METR's original paper. Each task in a suite gets a human time: how long a skilled engineer takes to do it, using the geometric mean of successful timed attempts. Each model then attempts the tasks several times. The 50% time horizon is the human time at which the model's success probability crosses one half.

The 50% horizon is not measured directly. It is read off a fitted curve. For each model j, the standard recipe fits a logistic regression, with log2(human time) as the only covariate:

logit p_j(t) = alpha_j - beta_j * log2(t)

Then you invert it: find the t where p_j(t) = 0.5. The intercept and slope are per model, and a 2025 METR research note moved to a shared slope across models. The new paper treats this shared-slope version as its baseline. Note that release dates are not in the fit. The exponential trend comes from plotting the fitted horizons against dates afterward.

The authors point out that this hides an assumption. A logistic regression on log time assumes that task difficulty for an AI grows linearly with the logarithm of human time. If that holds, every 10x step in human time costs the same amount of success probability. Nothing in the data collection guarantees it.

Relaxing the linearity assumption

The paper keeps the same data, which is 228 tasks in 79 task families and 26 models, and fits two richer statistical models. (The authors use "model" for the statistical fit and "AI" for the language models, a convention worth borrowing.)

Model 1 replaces the linear term with a monotone spline, so the link between human time and difficulty can bend.

Model 2 is an explanatory item-response theory (IRT) model. Item-response theory is the machinery behind standardized tests: it jointly estimates a respondent's ability and each item's difficulty from a grid of right and wrong answers. Here each task gets a latent difficulty. That latent value is regressed on human time through the spline, plus a task-family effect and a per-task residual. The spline can be read as a conversion function from human minutes to AI difficulty.

Model 2 also addresses a correlation problem. If runs were independent given the fitted probabilities, about 60% of a model's run sets on a task would be unanimous, either all successes or all failures. In the data, 83% are. Model 2 includes a correlation parameter that downweights task-model pairs with many redundant runs. The family effect replaces the heuristic square-root reweighting that METR used so large families did not dominate.

To compare fits, the authors use 5-fold cross-validation split by task family, so the held-out data comes from families the fit never saw. They score with proper scoring rules (log score, Brier score, and an elementary score at q = 0.5 and 0.8 that directly judges the horizon estimate), crossed with three weighting schemes. They report that their time-horizon estimates perform better under this suite than the baselines.

The flat region between 2 and 30 minutes

The most interesting result is the shape of the fitted conversion function. It is nearly flat from roughly 2 to 30 minutes of human time, and close to linear outside that range. In practical terms, tasks in the 2-30 minute band have similar AI difficulty even though their human times differ by 15x. Past 30 minutes, difficulty climbs with time again.

The consequence is that equal multipliers are not equal progress. The authors' example: a horizon jump from 3 minutes to 30 minutes is much easier than one from 30 minutes to 5 hours, despite both being 10x. Under the linear assumption they look identical on the plot. The paper's conditional success trajectory plot suggests a move from 4 to 15 minutes may be quite small in difficulty terms.

One plausible reading, offered here as interpretation rather than a finding from the paper, is that short tasks mostly test whether a model can follow instructions and call tools without slipping, while longer tasks add planning, recovery from errors, and context management. If the cost of those capabilities does not scale smoothly with duration, a single slope will blur them together.

The authors do not challenge the exponential growth of the plotted points. Their claim is about interpretation: the same visual slope can hide different amounts of capability gain depending on where it occurs.

What this changes for engineers

First, treat a time horizon as an estimate with a modeling choice behind it. If you cite "the model can do 2-hour tasks," ask which fit produced it and whether the tasks near that length are dense in the benchmark. A horizon estimated where tasks are sparse leans heavily on the assumed functional form.

Second, if you build your own agent evals and want to report capability in human-time units, the method transfers. The authors note that it applies to any IRT setup anchored on an external scale. METR's write-up of the original work is a good starting point for the task-timing protocol. About 29% of the tasks in the new analysis rely on an expert estimate of human time because timed attempts were missing or invalid, so measurement error in the x-axis is a real concern.

Third, use the diagnostic plots. The paper recommends reading time horizons together with the conversion plot and the conditional success trajectory, especially as benchmarks add longer tasks. A suite with few tasks beyond several hours cannot say much about progress there.

This is a preprint, a day old at the time of writing, and I have not reproduced the fits. The analysis covers one task suite, and the authors reserve the question of how these results extend to other domains and longer tasks for discussion. The spline's flat region is estimated from a finite number of tasks, so its exact boundaries carry uncertainty.

Conclusion

The time-horizon plot is useful because it puts capability in units people understand. The new paper's contribution is to show that the conversion from those units to difficulty is not a straight line, and to supply tools to check it. For practitioners the takeaway is simple: use the horizon number, but check which part of the human-time scale it sits in before comparing two jumps.

📰 Read the original article on Dev.to AI

Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.