Refuse the Percentage Until the Run Envelope Is Frozen
You opened the morning export and almost pasted it into a team note. Twelve coding tasks. Nine green. The sentence in your head was short: 75 percent, overnight, on a free server. Then the raw log spoiled it. Task 4 had
You opened the morning export and almost pasted it into a team note. Twelve coding tasks. Nine green. The sentence in your head was short: 75 percent, overnight, on a free server.
Then the raw log spoiled it. Task 4 had a cold start and a second attempt. Task 7 died on a timeout before any patch existed. Task 11 looked green, but the token-budget cell was empty. The percentage is a caption. It is not a result.
This walkthrough shows you how to freeze a run envelope, log host confounders, and block a headline rate when the envelope is incomplete. You get a dataset card, three metrics, a control table, and a small Python gate. Nothing here ranks a model. The script was not executed against a live host for this draft.
Name the question before the run
A citeable score answers one sentence you wrote first. Example: on this frozen task list, with this hidden-test hash, this timeout, and this token budget, how many tasks passed on the first completed attempt?
That sentence does not mean a product is better. It does not mean a free host is production-ready. If you cannot say the question without adjectives, stop. You are about to publish a demo.
Write the question into the same directory as the logs. Future you will try to widen it. The file is the brake.
1. Freeze the dataset card
Rerunability beats size. Eight to twenty tasks you can repeat will teach you more than a hundred you run once.
Give every task the same fields, and pin repo_revision to a commit id rather than a branch. Hash the prompt and the hidden tests. Do not commit those tests beside the agent.
Required card fields:
task_idrepo_revisionprompt_sha256hidden_tests_sha256allowed_filestimeout_stoken_budget
This row is a schema example. It is not a measured outcome.
task_id: calc-parse-017
repo_revision: 9f3c1ab
prompt_sha256: 6d2ec1aa
hidden_tests_sha256: a91b0440
allowed_files: ["src/parse.py"]
timeout_s: 180
token_budget: 12000
If two runs differ on any field, they are different experiments. Keep them in different files. Do not pool them later because the pass rates look similar.
2. Keep three metrics, not one
You will be tempted to publish a single rate. Don't. Store three, and print them together.
-
first_complete_pass_rateis the share of tasks whose first completed attempt passed the hidden tests. -
retry_rateis the share of tasks with a second attempt, a cold start, or a transport error. -
truncation_rateis the share of first attempts that stopped on timeout or token budget, not on a finished patch.
Add patch_apply_fail_rate when the harness cannot apply the diff. That failure belongs to you, not the agent. A task with an apply error must not sit in the denominator as a model fail unless you label the loss as harness loss.
Unequal budgets still must not be averaged. That older rule still holds. What changes here is the host: a retry measures the path, so it stays next to the pass rate.
3. Treat the server as a treatment
Shared compute changes the experiment even when the prompt file does not change. Queue wait, a cold start, a concurrency cap, or a silent retry can move the green count. Leave those out, and readers will blame the agent for a scheduler.
Disclosure: This article was prepared as part of MonkeyCode's product outreach.
Two availability claims are operator-supplied for this draft: free model access, and a free server option. You may use that pair as the place where you execute this method. You may not turn the product name into a metric.
This draft states no token quota, hardware shape, free-window length, or permanence promise. Those details were not checked against a vendor status page here. Write the class you actually used into access_class and server_class.
| Field | What it holds still | If you skip it |
|---|---|---|
| access_class | Free and paid paths may queue differently | Do not compare the runs |
| server_class | Shared CPU is part of the treatment | Say hosted, not model-only |
| concurrency | Parallel jobs share one cap | Headline rate is invalid |
| timeout_s | Slow correct patches get cut | Split truncation out |
| token_budget | Budget is part of the task | Refuse any average |
| attempt_policy | Retries hide transport noise | Keep retry_rate in the sentence |
4. Run a citeability gate
You want a script that refuses a percentage when the envelope is thin. The program below is a proposal, and it was not run on a live agent for this article. It does not open a socket. Point it at JSONL you already trust.
#!/usr/bin/env python3
"""Citeability gate. Proposal only. No network calls."""
import json
import sys
from pathlib import Path
REQUIRED = [
"task_id", "prompt_sha256", "hidden_tests_sha256",
"access_class", "server_class", "timeout_s", "token_budget",
"attempt", "cold_start", "queue_wait_ms", "stop_reason",
"patch_sha256",
]
def load_jsonl(path):
rows = []
for line in Path(path).read_text().splitlines():
if line.strip():
rows.append(json.loads(line))
return rows
def gate(rows):
missing = []
for i, row in enumerate(rows):
for key in REQUIRED:
if row.get(key) in (None, ""):
missing.append({"row": i, "task_id": row.get("task_id"), "field": key})
if missing:
return {"citeable": False, "reason": "incomplete_envelope", "missing": missing}
first = {}
for row in sorted(rows, key=lambda r: (r["task_id"], r["attempt"])):
first.setdefault(row["task_id"], row)
n = len(first)
passed = sum(1 for r in first.values() if r["stop_reason"] == "pass")
retried = {
r["task_id"] for r in rows
if r["attempt"] > 1 or r["cold_start"] or r["stop_reason"] == "transport"
}
truncated = sum(1 for r in first.values() if r["stop_reason"] in {"timeout", "budget"})
return {
"citeable": True,
"n_tasks": n,
"first_complete_pass_rate": round(passed / n, 4),
"retry_rate": round(len(retried) / n, 4),
"truncation_rate": round(truncated / n, 4),
"note": "Single envelope only. Not a product ranking.",
}
if __name__ == "__main__":
report = gate(load_jsonl(sys.argv[1]))
json.dump(report, sys.stdout, indent=2)
sys.stdout.write("\n")
sys.exit(0 if report["citeable"] else 2)
Save the next object as ok.jsonl. Every required field is present, and cold_start is false. The row is synthetic.
{"task_id":"calc-parse-017","prompt_sha256":"6d2ec1aa","hidden_tests_sha256":"a91b0440","access_class":"free_model","server_class":"free_server","timeout_s":180,"token_budget":12000,"attempt":1,"cold_start":false,"queue_wait_ms":840,"stop_reason":"pass","patch_sha256":"bb10"}
Save the following object as gap.jsonl. Queue wait is absent. That hole is the failure you want the gate to catch.
{"task_id":"calc-parse-017","prompt_sha256":"6d2ec1aa","hidden_tests_sha256":"a91b0440","access_class":"free_model","server_class":"free_server","timeout_s":180,"token_budget":12000,"attempt":1,"cold_start":false,"stop_reason":"pass","patch_sha256":"bb10"}
Run both commands. The second should exit 2.
python3 cite_gate.py ok.jsonl
echo "ok_exit=$?"
python3 cite_gate.py gap.jsonl
echo "gap_exit=$?"
If your local run of ok.jsonl matches the proposal, the shape looks like this. Treat it as the expected shape of the function, not as a captured host measurement.
citeable: true
n_tasks: 1
first_complete_pass_rate: 1.0
retry_rate: 0.0
truncation_rate: 0.0
Read the exit code before the pretty JSON. Exit 2 means notes, not a rate. A blank queue_wait_ms is enough. A dashboard screenshot is not a substitute.
5. Decide what sentence you may ship
| Gate result | You may write | You may not write |
|---|---|---|
| citeable false | Envelope incomplete. | Any percentage |
| citeable and retry_rate above 0.2 | Pass rate on a retry-heavy path. | The agent scores P. |
| citeable and truncation_rate above 0 | Some tasks stopped on budget or timeout. | Those fails were bad patches. |
| one access class and one server class | On this frozen envelope, first-complete pass rate was P (n=N), retry rate R. | Better than another product. |
| mixed server classes in one file | Split the file and rerun. | A pooled average |
The allowed sentence is dull on purpose. Dull is how you keep it out of a launch post. Drop n, drop the retry rate, or add a ranking verb, and the same arithmetic becomes marketing.
What the number still cannot prove
You hashed one prompt and one hidden suite. You did not sample the world's repositories. A second night on the same free server can change queue wait. Until you repeat the envelope, you have one draw.
Hidden tests stop you from grading prose. They do not detect training leakage. This gate does not check that. If the tasks are public puzzles, say so next to the rate, or do not cite the rate as generalization.
Do not treat a free server as an SLO oracle. Shared capacity is a fair place to debug the method. It is a bad place to promise latency.
Who should not use this gate
Skip it when you are showing a UI and you will not quote a rate. A tour does not need a denominator.
Skip it when you have no hidden tests. A model agreeing with its own summary is not an oracle. Build the tests first, then the gate.
Do not aim the agent at machines, accounts, or repos you are not allowed to modify. Keep the harness on fixtures you own. The method measures patches you can revert. It is not a scanning workflow.
If your decision is a paid capacity purchase, this design is the wrong instrument. It will not tell you how the same agent behaves once the queue disappears.
Put the gate in the commit
Drop cite_gate.py next to the task card. Run it on last week's JSONL before a percentage reaches a README or a slide. If the gate exits 2, publish the missing-field list instead. That list is the useful artifact.
If free model access plus a free server is where you can afford to repeat the same card, use that access to rerun the envelope. Keep the product name out of the title. One lab note with n, a retry rate, and a frozen hash will outlive a louder score.
Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.