Dev.to AI πŸ€– Ai πŸ‘ 0 πŸ“– 5 min read

Beyond Leaderboard Illusions: Benchmarking Multi-Turn Agentic Feedback Loops in Autonomous Software Engineering with TFD-Bench

Beyond Leaderboard Illusions: Benchmarking Multi-Turn Agentic Feedback Loops in Autonomous Software Engineering with TFD-Bench Submission for the Kaggle Benchmarking Challenge on DEV Author: Raja Rajak (@rajrajak99) K

Beyond Leaderboard Illusions: Benchmarking Multi-Turn Agentic Feedback Loops in Autonomous Software Engineering with TFD-Bench

Submission for the Kaggle Benchmarking Challenge on DEV

Author: Raja Rajak (@rajrajak99)

Kaggle Benchmark Dataset: Gemma 4 TFD Agentic Trajectories

Kaggle Evaluation Notebook: TFD-Bench on Kaggle

GitHub Repository: github.com/rajrajak99/gemma4-tfd-agent

⚑ Executive Summary & TL;DR

Standard AI coding benchmarksβ€”like HumanEval, MBPP, and static LeetCode-style puzzlesβ€”suffer from a fatal blindspot: they test one-shot code synthesis in isolation, completely ignoring how software engineering actually happens in the real world. Real software development is not a single prompt-response transaction; it is stateful, iterative, hypothesis-driven, and governed by test feedback.

When deployed inside live repositories, state-of-the-art LLMs suffer from what we call the "Illusion of One-Shot Competence": they generate elegant, syntactically clean code patches that fail 64.2% of real-world regression test suites.

To measure what actually matters, we designed TFD-Bench (Test-Feedback-Driven Benchmark): an evaluation suite of 50 multi-turn debugging tasks across 10 Python error classes adapted from SWE-bench Lite. We benchmarked 5 frontier model configurations across Pass@1 Resolve Rate, AST Tool Calling Validity, Context Token Consumption, and Loop Stalling Frequency.

Our key finding? Closed-loop test execution feedback acts as a massive reasoning accelerator: enforcing autonomous test reproduction before code generation jumps issue resolution from 29.5% to 45.6% while slashing context token waste by 49.5%.

🎯 What task(s) did you run?

1. The Core Problem with Static Coding Benchmarks

Current LLM evaluation often relies on single-turn benchmarks where an LLM is given a docstring and generates a self-contained function. This fails to evaluate four indispensable developer capabilities:

  1. Codebase Exploration: Can the model navigate multi-file repositories using symbol search and selective reading without exceeding context limits?
  2. Defect Reproduction: Can the model synthesize an isolated, minimal test script that independently confirms the bug (exit_code != 0) before touching existing code?
  3. Surgical Patching: Can the model apply targeted line replacements rather than rewriting entire files?
  4. Automated Verification: Can the model execute its reproducer, parse terminal tracebacks, and self-correct if the test fails?

2. The TFD-Bench Protocol

TFD-Bench tests models across a strict, 5-phase closed-loop cycle:

  • Phase 1: Discovery: Codebase localization via search_code and view_file.
  • Phase 2: Test-First: Minimal standalone reproduction script synthesis (create_reproducer) asserting pre-fix failure (exit_code != 0).
  • Phase 3: Reasoning & Repair: Surgical patch generation targeting minimal line changes (edit_file_replace).
  • Phase 4: Verification: Sandbox execution re-running the reproducer to assert resolution (exit_code == 0).
  • Phase 5: Self-Correction Loop: If verification fails, traceback parsing and automated retry (max 3 cycles).

3. Benchmark Task Taxonomy (50 Real-World Repository Tasks)

The dataset contains 50 gold-standard, repository-grade tasks evenly distributed across 10 Python bug categories (5 curated tasks per class):

Bug Category Real-World Scenario Primary Failure Mode in LLMs
ZeroDivisionError Metrics / Normalization calculations when denominator is 0 Missed edge-case guard clauses
IndexError Stream parsing / Token chunking on empty sequences Out-of-bounds slicing assumptions
KeyError Nested JSON schema validation & configuration parsing Unhandled missing dictionary keys
TypeError Dynamic type conversions & nullable payload inputs String/None concatenation bugs
AttributeError Method calls on optional/uninitialized object instances NoneType member access
FileNotFoundError Relative path resolution across different execution roots Working directory path drift
ValueError Parsing non-standard string formats / timestamps Unchecked input validation
RecursionError Circular tree traversals & recursive graph resolvers Missing base termination conditions
BoundaryCondition Off-by-one pagination & sliding window edge cases Inclusive vs exclusive range mistakes
ResourceLeak Unclosed socket/file descriptors in exception blocks Missing context managers (with statements)

πŸ€– Which models did you run it against?

To rigorously assess open vs. closed models and the impact of specialized fine-tuning, we evaluated 5 representative models across different scales and paradigms:

  1. TFD-Agent Google Gemma 4 31B + QLoRA: Google's Gemma 4 31B fine-tuned with 4-bit QLoRA targeting all linear projection layers (q, k, v, o, gate, up, down_proj) specifically conditioned on the Test-Feedback-Driven protocol.
  2. Google Gemma 4 31B (Standard ReAct): The identical base weights of Gemma 4 31B deployed in a vanilla ReAct prompt loop without TFD conditioning.
  3. Meta Llama 3.1 70B Instruct (SWE-agent Framework): Frontier open-weights generalist model evaluated via the established Princeton SWE-agent bash interface.
  4. Qwen 2.5 Coder 32B Instruct (ReAct): A state-of-the-art open code-specialized model known for high HumanEval scores.
  5. OpenAI GPT-4o mini (ReAct Baseline): Proprietary frontier lightweight model representing commercial API agent backends.

πŸ’‘ What are the main insights?

1. The Benchmark Leaderboard

Model Architecture Parameter Scale Pass@1 Resolve Rate (%) Tool Syntax Validity (%) Mean Tokens / Solved Task Loop Stalling Rate (%) Avg Turns to Fix
TFD-Agent Gemma 4 31B 31B 45.6% 96.9% 19,400 2.1% 7.2
Llama 3.1 70B (SWE-agent) 70B 42.4% 90.2% 32,100 8.5% 9.8
Qwen 2.5 Coder 32B (ReAct) 32B 41.0% 89.8% 29,800 9.6% 10.5
GPT-4o mini (ReAct) Frontier API 38.2% 87.6% 28,400 11.4% 11.2
Gemma 4 31B (Vanilla ReAct) 31B 29.5% 81.5% 38,400 18.2% 14.8

2. Discovery #1: The "Illusion of One-Shot Competence" (Silent Regressions)

When models operate without mandatory test reproduction, they exhibit an alarming failure mode:

  • In 64.2% of failed attempts, the vanilla models generated a patch that looked completely correct to human reviewers at a glance, but either failed to handle zero-length inputs, introduced a secondary exception, or broke downstream calling functions.
  • By forcing the model to write a reproducer FIRST (assert func(...) == ...) and confirm it fails with exit_code != 0, the agent grounds its reasoning in execution reality. This single procedural constraint boosted solve rates by +16.1 percentage points across the board.

3. Discovery #2: AST Tool Syntax Collapse Under Turn Drift

One of the most surprising findings was how tool invocation quality degrades over multi-turn conversations:

  • In turns 1 through 4, models achieve ~92% valid JSON tool calls.
  • Past turn 7, vanilla models suffer a 3.8x spike in tool syntax errors (malformed JSON strings, trailing commas, hallucinated tool parameters like path instead of file_path).
  • Once an error occurs, vanilla models often enter "hallucination cascades", repeating invalid tool arguments indefinitely.
  • Fine-tuning with turn-level supervision (TFD-Agent) kept tool validity above 96.9%, eliminating catastrophic loop stalling (dropping from 18.2% to 2.1%).

4. Discovery #3: Test Feedback as a Reasoning & Token Compression Engine

Many developers assume adding automated test runs increases inference costs. Our benchmark proved the exact opposite:

  • Standard ReAct Loop: 38,400 tokens per solved task (model wanders aimlessly searching files and trying random variations).
  • TFD-Agent Closed Loop: 19,400 tokens per solved task (a 49.5% context reduction).
  • Why? A concrete traceback with exit_code: 1 and line numbers acts as a semantic anchor. The model doesn't need to guess where the fault is; the Python runtime tells it directly, pruning redundant exploration paths.

5. Discovery #4: Fine-Tuning Beats Scale in Agentic Loops

Notice that TFD-Agent (31B) outperformed Llama 3.1 70B (45.6% vs 42.4%) despite having less than half the parameters:

  • General-purpose 70B models have superior world knowledge, but they are not conditioned to adhere to the strict protocol of Locate -> Reproduce -> Surgical Edit -> Verify.
  • When a 31B model is aligned to treat terminal output as ground-truth feedback, it achieves state-of-the-art software engineering competence on consumer/edge workstations.

πŸ” Case Study: Anatomy of an Autonomous Fix

Here is a real example from task tfd_zerodivisionerror_1 demonstrating the power of the loop:

Step 1: Locating the Fault


json
{"tool": "search_code", "arguments": {"query": "def calculate_precision_recall"}}
πŸ“° Read the original article on Dev.to AI

Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes β€” full credit and traffic to the original publisher.