Green Property Checks Must Not Refill a Flake Freeze Token
A flake freeze is a single-spend token. A later property pass does not refill its replay budget, and a fixture match does not move its expiry. Those three results answer different questions. Folding them into one green b
A flake freeze is a single-spend token. A later property pass does not refill its replay budget, and a fixture match does not move its expiry. Those three results answer different questions. Folding them into one green bit is how a temporary waiver becomes a permanent merge rule.
Agent patches make that collapse easy. A diff can preserve every locked output and still break an invariant the fixture never named. A diff can trip one known flake and still be the wrong change to merge. Status codes hide the split. This note specifies a write policy for the freeze ledger: automated checks may spend a token, and they may not mint or refill one.
The program below is an illustrative specification with a fixed clock of 2026-10-09T00:00:00Z. It is not a report of production failure rates. The outcomes listed with it are the ones the example is written to produce.
Keep three answers in three fields
Property checks state whether the diff preserves invariants you can evaluate without the golden file. Fixture locks state whether observed outputs still hash to the locked digest. A freeze token states whether a reviewer already classified this failure signature as non-blocking, and whether that classification still has budget.
Each check writes only its own field. A property pass stores property=pass. It does not increment remaining_replays. A fixture match stores fixture=match. It does not assign a new expires_at. Only a mint command creates a token, and that command requires a reviewer id the patch author cannot supply.
Abstain is not a pass. If the invariant module fails to import, the verdict is abstain and admission stops. If the fixture file is absent, the verdict is missing, not match. Lookup order is not the claim here. The claim is which fields a check may write once a token row is already in hand.
Read the matrix before you automate it
Rows are inputs. The last column is the only allowed action. write is true only when the ledger row itself must change.
| Property | Fixture | Token | Action | Write row |
|---|---|---|---|---|
| fail | any | any | reject, do not read the token | no |
| abstain | any | any | quarantine, do not mint | no |
| pass | missing | any | quarantine until a fixture lock exists | no |
| pass | mismatch | none or expired | reject, no inherited budget | only to mark expiry if a row exists and is spent out |
| pass | mismatch | live, budget > 0 | spend one, hold, do not merge | yes, budget only |
| pass | match | any | admit, leave the token untouched | no |
The last row is the non-refill rule. A match may admit this patch. It must not extend some other token that shares the test id. Two signatures are different rows. Hash the test id, the normalized message, and the fixture id. Do not put the patch body in that key.
A regenerated diff is a new patch_id, but it can still spend a live token if it reproduces the same signature. It still cannot refill that token. Spending and refilling are different writes. Keeping them apart is the whole gate.
Normalize the signature, then mint
Unstable text defeats the budget. Timestamps and random ids inside assertion messages create a new key on every run, so the ledger never spends the row you meant to cap.
- Strip timestamps and generated ids from the failure message before hashing.
- Build
failure_signaturefrom test id, normalized message, and fixture id. - Reject the mint if any of those parts is empty.
- Store
reviewer,reason_code,remaining_replays, andexpires_atin UTC on the ledger clock. - Refuse mint from the process that wrote the patch. A separate command must supply the reviewer id.
A local normalization check, using a fixed sample rather than a live log:
python3 - <<'PY'
import hashlib, re
msg = "race on list at 2026-10-09T11:02:03Z id=91ab"
norm = re.sub(r"\d{4}-\d{2}-\d{2}T[\d:.]+Z", "<ts>", msg)
norm = re.sub(r"id=[0-9a-f]+", "id=<id>", norm)
print(norm)
print(hashlib.sha256(b"t_order|" + norm.encode() + b"|fx_order").hexdigest()[:16])
PY
The specified normalized message is race on list at <ts> id=<id>. If your real logs do not reduce to a stable string under the same substitutions, do not mint. Add a substitution, or drop the freeze. A token on a moving key is an unbounded waiver with a hash in front of it.
Apply the write policy in order
- Hash the diff and store
patch_id. Do not call a model in this step. - Run property checks in-process. Record
pass,fail, orabstain. Onfailorabstain, stop. Do not open the token store. - Replay locked fixtures. Record
match,mismatch, ormissing. Onmissing, stop and open a lock task. - On
mismatch, derivefailure_signatureand load the token for that signature only. - If the row is missing, or
expires_atis not after the ledger clock, reject. Do not auto-mint. - If
remaining_replaysis zero, mark the row expired and reject. That write records exhaustion. It does not grant a new budget. - If budget remains, decrement it by one with a compare-and-swap, then hold. A hold is not a merge.
- On
passplusmatch, admit and setwrite=false. Do not persist the token object, even when every field looks unchanged. - Mint only through the separate command above. Automation may spend. It may not create, and it may not refill after a pass.
Step 8 is the one teams skip when a green suite feels like evidence that the flake is gone. Green evidence can admit this patch. It is not evidence that a future miss should receive a larger budget. Persisting the returned object when write is false is a defect, because a later edit inside decide could refill fields and the save would commit them.
Specification you can run
Save the following as freeze_token_gate.py. It uses a fixed clock so the result does not depend on the runner's time. It does not open a socket. Treat it as a specification of the write bit, not as a measured CI report.
import hashlib
import json
from datetime import datetime, timedelta, timezone
FIXED_NOW = datetime(2026, 10, 9, tzinfo=timezone.utc)
def signature(test_id, message, fixture_id):
raw = f"{test_id}|{message.strip()}|{fixture_id}".encode()
return hashlib.sha256(raw).hexdigest()[:16]
def mint(test_id, message, fixture_id, reviewer, budget, hours, now=FIXED_NOW):
if not reviewer or budget < 1 or hours < 1:
raise ValueError("mint requires reviewer, budget >= 1, hours >= 1")
return {
"signature": signature(test_id, message, fixture_id),
"reviewer": reviewer,
"reason_code": "known-order-flake",
"remaining_replays": budget,
"expires_at": now + timedelta(hours=hours),
}
def decide(prop, fixture, token, now=FIXED_NOW):
if prop == "fail":
return "reject", token, False
if prop == "abstain" or fixture == "missing":
return "quarantine", token, False
if prop == "pass" and fixture == "match":
return "admit", token, False
if fixture != "mismatch" or token is None:
return "reject", token, False
if token["expires_at"] <= now or token["remaining_replays"] <= 0:
expired = dict(token)
expired["remaining_replays"] = 0
return "reject", expired, True
spent = dict(token)
spent["remaining_replays"] = token["remaining_replays"] - 1
return "hold", spent, True
def demo():
token = mint("t_order", "race on list", "fx_order", "rev-14", 2, 24)
cases = [
("fail", "mismatch", token),
("pass", "match", token),
("pass", "mismatch", token),
("abstain", "match", token),
]
for prop, fixture, tok in cases:
action, after, write = decide(prop, fixture, tok)
print(json.dumps({
"property": prop,
"fixture": fixture,
"action": action,
"write": write,
"budget_after": None if after is None else after["remaining_replays"],
"expiry_unchanged": after is not None and after["expires_at"] == token["expires_at"],
}))
if __name__ == "__main__":
demo()
Specified results if you run the file as written:
-
failreturnsreject,write=false, budget still2. -
passandmatchreturnadmit,write=false, budget still2, same expiry. -
passandmismatchreturnhold,write=true, budget1, same expiry. -
abstainreturnsquarantine,write=false, budget still2.
Pass the spent object back in for a second hold and the next call must reject with budget 0. That sequence is the replay budget. It is not a retry loop, and it is not a reason to call a model. The expiry timestamp stays on the original mint.
Lock the non-refill rule in tests/test_freeze_token_gate.py:
from freeze_token_gate import FIXED_NOW, decide, mint
def test_admit_does_not_refill_or_extend():
token = mint("t_order", "race on list", "fx_order", "rev-14", 1, 4)
action, after, write = decide("pass", "match", token, FIXED_NOW)
assert action == "admit"
assert write is False
assert after["remaining_replays"] == 1
assert after["expires_at"] == token["expires_at"]
def test_spend_does_not_move_expiry_and_second_spend_rejects():
token = mint("t_order", "race on list", "fx_order", "rev-14", 1, 4)
action, spent, write = decide("pass", "mismatch", token, FIXED_NOW)
assert action == "hold"
assert write is True
assert spent["remaining_replays"] == 0
assert spent["expires_at"] == token["expires_at"]
action2, expired, write2 = decide("pass", "mismatch", spent, FIXED_NOW)
assert action2 == "reject"
assert write2 is True
assert expired["remaining_replays"] == 0
assert expired["expires_at"] == token["expires_at"]
python3 freeze_token_gate.py
python3 -m pytest -q tests/test_freeze_token_gate.py
If test_admit_does_not_refill_or_extend fails, the admit path is writing the freeze row. Stop and remove that write before you add any network client to the job. A green suite that also refreshes expires_at is not a stronger suite. It is a waiver that grows on success.
Where a free draft is allowed to sit
Steps 1 through 8 are local. A model call is eligible only after the ledger returns reject for a reason code you have already marked regenerable, and only to produce a new diff with a new patch_id. A hold is not eligible. The miss is already classified, and another draft must not spend or refill that row.
Disclosure: This article was prepared as part of MonkeyCode's product outreach.
MonkeyCode's free model access fits as a draft source for that new diff. Its free server option fits as an isolated replay host, so the failure signature is taken from a clean run rather than from leftover state on a shared runner. Neither service is the freeze authority. The mint command still requires a reviewer id that the draft process cannot supply. Those two availability claims are operator-stated options. They are not a quota, a hardware shape, a latency bound, or a promise that the free path remains. If that path is down, run the same ledger on another replay host. Do not skip the token rules because one endpoint is unreachable.
Eligibility for the draft call is a policy list, not a measured win rate:
- Regenerable reason codes:
format,import-order,unused-symbol. - Not regenerable: invariant failure, missing fixture, abstain, expired token, exhausted budget.
- Never on a hold, and never as a substitute for mint.
Keep the property module and the fixture files out of the draft's editable paths. A model that "fixes" a flake by editing the invariant, or by rewriting the golden file, is manufacturing a pass. Treat empty draft output as abstain, not as a signal to retry the same call. Then leave the token row alone.
If your review log can already store a reviewer id apart from the diff author, run the non-refill test on one flaky signature before you attach any draft client.
Limitations
Signature collision is the main hole. A real regression can emit the same normalized message as a known flake. The ledger will then spend the flake token and hold a patch that should have been rejected. Cap the budget low, and require the reason code to name the flake mechanism, not just the assertion text.
Clock skew is the second hole. Compare expires_at to the ledger clock in UTC. A runner clock can make a dead token look live. That is why the example pins FIXED_NOW, and why production code should not trust log timestamps as the authority.
The in-memory store is not safe under two runners. Both can read budget 1 and both can hold unless the decrement is atomic. That bug looks like a flake and is not one. Use a compare-and-swap on (signature, remaining_replays) before you trust a hold.
This gate does not prove the patch is correct. It proves that three answers were not allowed to overwrite each other. Semantic review stays outside the table. A passing property check is only as strong as the invariants you wrote down. If those invariants are restatements of the fixture file, you have one check with two names.
Free-tier drafts add a further limit. The draft can be empty, off-topic, or a test edit that forces a fixture match. None of those outcomes may set write=true on a freeze row. Availability can also change. Build the replay so a missing draft host degrades to "no regeneration," not to "skip the ledger."
Who should skip it
Skip the ledger if you cannot produce a stable signature. A freeze on a key that changes every run is an unbounded waiver with extra columns.
Skip it if the property command and the fixture command are the same process. You would be storing one answer in two fields and calling the split a control. Split the commands first, or do not pretend the table is doing work.
Skip it on an incident hotfix that must ship before anyone can mint. Use the incident path. Do not mint after the fact to launder that hotfix into the freeze table. A backdated token hides the bypass from the same audit this design is meant to support.
Skip it if you want a model to decide that a failure is flaky. That decision is the mint. Putting it on the process that wrote the patch removes the separation the write bit exists to keep.
What to retain for audit
Keep five fields on every hold or reject: patch_id, property verdict, fixture verdict, token signature or none, and the write bit. Those fields are enough to query whether an admit ever persisted expires_at or remaining_replays. If a later change refills budget on the green path, the query fails without a narrative review.
-- illustrative audit, not a vendor-specific dialect
select patch_id, property, fixture, token_signature, write_bit
from patch_gate_log
where action = 'admit' and write_bit = true;
An empty result is the expected steady state for the non-refill rule. Any row that query returns is a gate defect, even if the suite was green that day.
The operational conclusion is narrow. Spend a freeze token once. Expire it on budget or on the ledger clock. Refuse every automated path, including a free draft, that tries to refill it. Property checks and fixture locks keep their own columns, and a green result is not a mint.
Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes โ full credit and traffic to the original publisher.