Dev.to AI 🤖 Ai 👁 0 📖 4 min read

Tuning a Local Qwen Coding Agent: What Changed, What Still Fails

On eight development coding cases repeated three times, a local Qwen3.8-27B agent went from 5/24 to 23/24 functional and delivered successes across sequential output-budget and reasoning rounds. That is promising develop

On eight development coding cases repeated three times, a local Qwen3.8-27B agent went from 5/24 to 23/24 functional and delivered successes across sequential output-budget and reasoning rounds. That is promising development evidence. It does not establish generalization or show that xhigh alone caused the difference.

The next experiment was less encouraging: lowering temperature to 0.8 on two selected difficult cases preserved 5/6 functional successes but reduced delivered successes to 4/6. The reused temperature-1.0 controls scored 5/6 on both metrics.

This is a follow-up to my local OpenCode experiments. The full article and public records provide the detailed methods and inclusion rules. This post is a self-contained summary of the development findings and their limits.

Response space and reasoning

The setup used Qwen3.8-27B Q4_K_M on one RTX 3090, llama.cpp b11146-7fe450e19, CUDA 12.8, and OpenCode 2.0.20. OpenCode operated on the workspace and ran tools; llama-server supplied model responses. These agentic-v2.0.1 results are separate from the earlier deployment and MTP studies.

Recorded profile Functional Delivered
Medium, 8,192 output tokens 5/24 5/24
Medium, 16,384 output tokens 14/24 14/24
Medium, 32,768 output tokens 20/24 20/24
Xhigh, 32,768 output tokens 23/24 23/24

Each row contains the same eight cases with three matched repeat labels. The output allowance is per response and includes reasoning. It is distinct from the 131,072-token context capacity and the whole-attempt limits of 1,800 seconds, 65,536 generated tokens, and 120 tool calls.

More response space coincided with more completed repairs, but these were sequential development rounds selected after feedback. Reused controls, correlated repeats, and changes in rest/recovery history prevent treating the table as a randomized causal study. The last row is not a universal coding success rate.

Correct patches and finished work are different

Functional success required solved phase-one and final grades and unchanged agent configuration. Solved grading included protected-file and starter-file checks. Python required independent acceptance/regression checks, public tests, and syntax checks. TypeScript required acceptance/regression checks, public tests, and project and consumer type checks.

Delivered success additionally required the expected turn count, normal completed termination, a nonempty final response after tools, and the last recognized agent test validation passing. TypeScript also required the last recognized typecheck to pass. The development cases each required one turn.

The targeted temperature round shows why both outcomes matter:

  • One ledger attempt passed both grades but reached the 1,800-second wall-time limit without a final response: functional success, failed delivery.
  • Another exited normally, reported passing agent tests, and supplied final text, but failed independent acceptance checks: neither success metric passed.

Passing the frozen checks is not proof of a defect-free patch. Final-text presence also does not establish that the handoff accurately describes the code.

Temperature 0.8 was not a better default in this test

Targeted profile Functional Delivered
Temperature 1.0, reused xhigh/32K controls 5/6 5/6
Temperature 0.8, new xhigh/32K attempts 5/6 4/6

These were ledger and pagination, three repeats each, selected after earlier failures. No fresh temperature-1.0 controls ran alongside the new attempts. The reused controls must not be counted again as new observations. The result supplies no basis for adopting 0.8 from this small targeted sample.

Confirmation did not run

The selected profile was xhigh reasoning, 32,768 output tokens, and temperature 1.0. A frozen confirmation protocol planned four fresh cases, three paired repeats, and two complete profiles: medium/8K versus xhigh/32K, both at temperature 1.0.

Safety preflight stopped launch because hardware readiness remained unresolved after a CPU machine-check report. A separate resource-ownership prerequisite was also unresolved. No defective component or hardware cause was established, and neither guard was bypassed.

Zero confirmation attempts ran. There are no confirmation quality measurements or complete pairs. Missing observations are not zero-percent success rates or model failures. Development results remain provisional, and repeated runs on four cases would still not create twelve independent cases.

Inspect the records

The public data README links the protocol, per-attempt JSON/CSV, aggregates, checksum file, and standalone checker. Save all eight files in one directory and run:

python3 check.py

Python 3.10 or later is sufficient. The checker uses no network or model and writes nothing. It verifies checksums, JSON/CSV parity, membership, inventory, grade/outcome formulas, totals, medians, and paired transitions. The four main cohorts plus six targeted attempts contain 102 unique primary observations; the broader inventory retains pilots, failures, infrastructure history, and unstarted slots separately.

This reproduces aggregation only. The package omits fixtures, prompts, patches, sessions, private tests, numeric seeds, and the inference/grading harness. It cannot rerun grading, prove source-record authenticity, or establish generalization.

The useful takeaway is to record correctness, delivery, and cost separately, keep failures visible, and lock a candidate before fresh confirmation. The development profile is worth testing on representative work; the unrun confirmation remains an open limitation.

Disclosure: AI assisted the preparation of this article from the recorded experiments and independently reviewed development export. The factual claims were checked against those records. This does not imply that a human independently reviewed every model-produced patch.

📰 Read the original article on Dev.to AI

Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.