Dev.to AI 🤖 Ai 👁 0 📖 2 min read

I Tested Whether AI Can Keep Up With a Changing Software Specification

This is a submission for the Kaggle Benchmarking Challenge. What I Benchmarked Software specifications rarely remain static. A team might replace password authentication with OAuth, tighten permissions, or in

This is a submission for the Kaggle Benchmarking Challenge.

What I Benchmarked

Software specifications rarely remain static. A team might replace password authentication with OAuth, tighten permissions, or introduce guest access while keeping guests read-only.

I built SpecDrift, a benchmark designed to test whether an AI model can reconstruct the current specification from a sequence of changing requirements.

My first task, Access & Identity, covers five scenarios:

  • Replacing email/password authentication with Google OAuth.
  • Restricting customer-data exports to administrators.
  • Replacing password authentication with email magic links.
  • Allowing guests to view public dashboard pages without editing them.
  • Renaming the owner role to administrator and transferring its permissions.

The task evaluates explicit final-state fields and changed-key identification rather than relying entirely on whether an explanation sounds convincing.

Models Tested

My initial baseline was google/gemini-3.7-flash.

I chose a structured-output evaluation because software teams need machine-readable specifications that downstream systems can use, not only plausible natural-language explanations.

I plan to expand the benchmark to additional available models and more requirement categories.

Findings

In my initial run, Gemini passed 16 of 18 assertions.

The two failures exposed an unexpected issue in the evaluator. The model returned email_magic_link for the final authentication method and read_only for guest permissions, but my regular expressions did not accept underscores.

This matters because an evaluation can incorrectly penalize a correct answer when its matching rules are too restrictive. The benchmark therefore needs testing at two levels: whether the model reconstructed the requirements correctly, and whether the evaluator correctly recognizes valid representations.

I updated the regex patterns to handle underscore-separated values and reran the task. The final reported result should reflect that rerun.

The broader question behind SpecDrift is whether models can reliably distinguish current requirements from obsolete ones, especially when a conversation contains several revisions. My next step is to extend the cases into billing, API integration, operations, and notifications, and compare more models using the same tests.

My Benchmark

https://www.kaggle.com/benchmarks/pratik222/specdrift-can-ai-track-changing-requirements

SpecDrift explores a practical challenge in AI-assisted software development: can a model keep the current truth when the specification changes?

kagglechallenge

📰 Read the original article on Dev.to AI

Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.