Beyond Leaderboard Illusions: Benchmarking Multi-Turn Agentic Feedback Loops in Autonomous Software Engineering with TFD-Bench
Beyond Leaderboard Illusions: Benchmarking Multi-Turn Agentic Feedback Loops in Autonomous Software Engineering with TFD-Bench Submission for the Kaggle Benchmarking Challenge on DEV Author: Raja Rajak (@rajrajak99) K
Beyond Leaderboard Illusions: Benchmarking Multi-Turn Agentic Feedback Loops in Autonomous Software Engineering with TFD-Bench
Submission for the Kaggle Benchmarking Challenge on DEV
Author: Raja Rajak (@rajrajak99)
Kaggle Benchmark Dataset: Gemma 4 TFD Agentic Trajectories
Kaggle Evaluation Notebook: TFD-Bench on Kaggle
GitHub Repository: github.com/rajrajak99/gemma4-tfd-agent
β‘ Executive Summary & TL;DR
Standard AI coding benchmarksβlike HumanEval, MBPP, and static LeetCode-style puzzlesβsuffer from a fatal blindspot: they test one-shot code synthesis in isolation, completely ignoring how software engineering actually happens in the real world. Real software development is not a single prompt-response transaction; it is stateful, iterative, hypothesis-driven, and governed by test feedback.
When deployed inside live repositories, state-of-the-art LLMs suffer from what we call the "Illusion of One-Shot Competence": they generate elegant, syntactically clean code patches that fail 64.2% of real-world regression test suites.
To measure what actually matters, we designed TFD-Bench (Test-Feedback-Driven Benchmark): an evaluation suite of 50 multi-turn debugging tasks across 10 Python error classes adapted from SWE-bench Lite. We benchmarked 5 frontier model configurations across Pass@1 Resolve Rate, AST Tool Calling Validity, Context Token Consumption, and Loop Stalling Frequency.
Our key finding? Closed-loop test execution feedback acts as a massive reasoning accelerator: enforcing autonomous test reproduction before code generation jumps issue resolution from 29.5% to 45.6% while slashing context token waste by 49.5%.
π― What task(s) did you run?
1. The Core Problem with Static Coding Benchmarks
Current LLM evaluation often relies on single-turn benchmarks where an LLM is given a docstring and generates a self-contained function. This fails to evaluate four indispensable developer capabilities:
- Codebase Exploration: Can the model navigate multi-file repositories using symbol search and selective reading without exceeding context limits?
-
Defect Reproduction: Can the model synthesize an isolated, minimal test script that independently confirms the bug (
exit_code != 0) before touching existing code? - Surgical Patching: Can the model apply targeted line replacements rather than rewriting entire files?
- Automated Verification: Can the model execute its reproducer, parse terminal tracebacks, and self-correct if the test fails?
2. The TFD-Bench Protocol
TFD-Bench tests models across a strict, 5-phase closed-loop cycle:
-
Phase 1: Discovery: Codebase localization via
search_codeandview_file. -
Phase 2: Test-First: Minimal standalone reproduction script synthesis (
create_reproducer) asserting pre-fix failure (exit_code != 0). -
Phase 3: Reasoning & Repair: Surgical patch generation targeting minimal line changes (
edit_file_replace). -
Phase 4: Verification: Sandbox execution re-running the reproducer to assert resolution (
exit_code == 0). - Phase 5: Self-Correction Loop: If verification fails, traceback parsing and automated retry (max 3 cycles).
3. Benchmark Task Taxonomy (50 Real-World Repository Tasks)
The dataset contains 50 gold-standard, repository-grade tasks evenly distributed across 10 Python bug categories (5 curated tasks per class):
| Bug Category | Real-World Scenario | Primary Failure Mode in LLMs |
|---|---|---|
| ZeroDivisionError | Metrics / Normalization calculations when denominator is 0 | Missed edge-case guard clauses |
| IndexError | Stream parsing / Token chunking on empty sequences | Out-of-bounds slicing assumptions |
| KeyError | Nested JSON schema validation & configuration parsing | Unhandled missing dictionary keys |
| TypeError | Dynamic type conversions & nullable payload inputs | String/None concatenation bugs |
| AttributeError | Method calls on optional/uninitialized object instances | NoneType member access |
| FileNotFoundError | Relative path resolution across different execution roots | Working directory path drift |
| ValueError | Parsing non-standard string formats / timestamps | Unchecked input validation |
| RecursionError | Circular tree traversals & recursive graph resolvers | Missing base termination conditions |
| BoundaryCondition | Off-by-one pagination & sliding window edge cases | Inclusive vs exclusive range mistakes |
| ResourceLeak | Unclosed socket/file descriptors in exception blocks | Missing context managers (with statements) |
π€ Which models did you run it against?
To rigorously assess open vs. closed models and the impact of specialized fine-tuning, we evaluated 5 representative models across different scales and paradigms:
-
TFD-Agent Google Gemma 4 31B + QLoRA:
Google's Gemma 4 31B fine-tuned with 4-bit QLoRA targeting all linear projection layers (
q, k, v, o, gate, up, down_proj) specifically conditioned on the Test-Feedback-Driven protocol. - Google Gemma 4 31B (Standard ReAct): The identical base weights of Gemma 4 31B deployed in a vanilla ReAct prompt loop without TFD conditioning.
- Meta Llama 3.1 70B Instruct (SWE-agent Framework): Frontier open-weights generalist model evaluated via the established Princeton SWE-agent bash interface.
- Qwen 2.5 Coder 32B Instruct (ReAct): A state-of-the-art open code-specialized model known for high HumanEval scores.
- OpenAI GPT-4o mini (ReAct Baseline): Proprietary frontier lightweight model representing commercial API agent backends.
π‘ What are the main insights?
1. The Benchmark Leaderboard
| Model Architecture | Parameter Scale | Pass@1 Resolve Rate (%) | Tool Syntax Validity (%) | Mean Tokens / Solved Task | Loop Stalling Rate (%) | Avg Turns to Fix |
|---|---|---|---|---|---|---|
| TFD-Agent Gemma 4 31B | 31B | 45.6% | 96.9% | 19,400 | 2.1% | 7.2 |
| Llama 3.1 70B (SWE-agent) | 70B | 42.4% | 90.2% | 32,100 | 8.5% | 9.8 |
| Qwen 2.5 Coder 32B (ReAct) | 32B | 41.0% | 89.8% | 29,800 | 9.6% | 10.5 |
| GPT-4o mini (ReAct) | Frontier API | 38.2% | 87.6% | 28,400 | 11.4% | 11.2 |
| Gemma 4 31B (Vanilla ReAct) | 31B | 29.5% | 81.5% | 38,400 | 18.2% | 14.8 |
2. Discovery #1: The "Illusion of One-Shot Competence" (Silent Regressions)
When models operate without mandatory test reproduction, they exhibit an alarming failure mode:
- In 64.2% of failed attempts, the vanilla models generated a patch that looked completely correct to human reviewers at a glance, but either failed to handle zero-length inputs, introduced a secondary exception, or broke downstream calling functions.
-
By forcing the model to write a reproducer FIRST (
assert func(...) == ...) and confirm it fails withexit_code != 0, the agent grounds its reasoning in execution reality. This single procedural constraint boosted solve rates by +16.1 percentage points across the board.
3. Discovery #2: AST Tool Syntax Collapse Under Turn Drift
One of the most surprising findings was how tool invocation quality degrades over multi-turn conversations:
- In turns 1 through 4, models achieve ~92% valid JSON tool calls.
-
Past turn 7, vanilla models suffer a 3.8x spike in tool syntax errors (malformed JSON strings, trailing commas, hallucinated tool parameters like
pathinstead offile_path). - Once an error occurs, vanilla models often enter "hallucination cascades", repeating invalid tool arguments indefinitely.
- Fine-tuning with turn-level supervision (TFD-Agent) kept tool validity above 96.9%, eliminating catastrophic loop stalling (dropping from 18.2% to 2.1%).
4. Discovery #3: Test Feedback as a Reasoning & Token Compression Engine
Many developers assume adding automated test runs increases inference costs. Our benchmark proved the exact opposite:
- Standard ReAct Loop: 38,400 tokens per solved task (model wanders aimlessly searching files and trying random variations).
- TFD-Agent Closed Loop: 19,400 tokens per solved task (a 49.5% context reduction).
-
Why? A concrete traceback with
exit_code: 1and line numbers acts as a semantic anchor. The model doesn't need to guess where the fault is; the Python runtime tells it directly, pruning redundant exploration paths.
5. Discovery #4: Fine-Tuning Beats Scale in Agentic Loops
Notice that TFD-Agent (31B) outperformed Llama 3.1 70B (45.6% vs 42.4%) despite having less than half the parameters:
- General-purpose 70B models have superior world knowledge, but they are not conditioned to adhere to the strict protocol of Locate -> Reproduce -> Surgical Edit -> Verify.
- When a 31B model is aligned to treat terminal output as ground-truth feedback, it achieves state-of-the-art software engineering competence on consumer/edge workstations.
π Case Study: Anatomy of an Autonomous Fix
Here is a real example from task tfd_zerodivisionerror_1 demonstrating the power of the loop:
Step 1: Locating the Fault
json
{"tool": "search_code", "arguments": {"query": "def calculate_precision_recall"}}
Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes β full credit and traffic to the original publisher.