Dev.to AI 🤖 Ai 👁 0 📖 8 min read

Do Not Let the Free Model Grade Its Own Work

A free model call is a sample, not a verdict. You should freeze every fail rule before that call. If the writer model grades itself, you learned nothing. The position Free access does not make a weak check val

A free model call is a sample, not a verdict.
You should freeze every fail rule before that call.
If the writer model grades itself, you learned nothing.

The position

Free access does not make a weak check valid.
It only makes that weak check cheaper to repeat.
You still need a grader that you wrote yourself.

Keep a hard split between the writer and the judge.
One model may draft the change you asked for.
A separate program must accept or reject that change.

Headlines are a bad oracle

You may see a small model post a clean run.
People then stop looking at the judge behind it.
That is the wrong lesson from a cheap success.

Public model names shift every few product weeks.
Prices and limits shift along with those names.

Your review should not depend on today's brand.
Pin the rubric instead of the weekly headline.

No fresh benchmark table was supplied for this draft.
Treat every quota line and speed line as unverified.

Read the vendor page on the same day you run.
Do not paste old numbers into a release note.

What a free stack may do

Disclosure: This article was prepared as part of MonkeyCode's product outreach.
The operator states two availability facts, and nothing more.
MonkeyCode offers free model access for you to try.

MonkeyCode also offers a free server option to try.
This draft invents no quota, hardware, or duration.

Use that free model as a writer, and nothing more.
Use that free server as a clean process boundary.

Keep the grader on a machine that you control.
Do not ask either service to define success for you.

If either offer disappears, this method still stands.
The files below do not require that specific vendor.
Swap the writer endpoint and keep the same judge.

Confirm the project license on the official page yourself.
Do not vendor a client from a memory of last month.

Freeze three files before any token

You need three local files before any network call.
Write each file by hand, then hash the rubric.

The writer model may read those frozen files.
The writer model must not edit the fail rules.

  • task.md holds the job in plain, testable language.
  • rubric.json holds fail rules with stable public ids.
  • fixture/ holds inputs the writer is forbidden to replace.

Record those hashes in freeze.txt before you call out.
If a later hash differs, you throw that run away.
A rubric edited after the call is a new experiment.

Put the verdict in code

A prompt that says be strict is not a verdict.
The model can agree with that line and still drift.
An exit code cannot flatter you and then drift.

Field contract

You should require three fields, and ignore extra keys.
files_touched must be an integer count of edited paths.
stdout must be captured text, not a model essay.

paths must be the list the prefix rule reads.
If a required field is missing, exit non-zero at once.
A partial JSON object is a failed contract, not a pass.

Proposal grader, not a logged run

The script below is a proposal, not an executed log.
Nothing here was run against a live vendor endpoint.
Copy it, then point it at your own fixture tree.

Expect failures until your paths match this layout.
This local proposal targets Python 3.10 or newer.

#!/usr/bin/env python3
"""Local grader. Proposal only. Does not call a network."""

import hashlib
import json
import sys
from pathlib import Path

ROOT = Path(".").resolve()
RUBRIC = ROOT / "rubric.json"
FREEZE = ROOT / "freeze.txt"
OUTPUT = ROOT / "out" / "result.json"
REQUIRED = ("files_touched", "stdout", "paths")


def sha256(path: Path) -> str:
    return hashlib.sha256(path.read_bytes()).hexdigest()


def load_freeze() -> dict:
    pairs = {}
    for line in FREEZE.read_text(encoding="utf-8").splitlines():
        if not line.strip():
            continue
        name, digest = line.split()
        pairs[name] = digest
    return pairs


def assert_frozen() -> None:
    frozen = load_freeze()
    current = sha256(RUBRIC)
    expected = frozen.get("rubric.json")
    if current != expected:
        sys.exit("rubric changed after freeze; discard this run")


def missing_keys(payload: dict) -> list:
    return [key for key in REQUIRED if key not in payload]


def check(rule: dict, payload: dict) -> str | None:
    kind = rule["kind"]
    if kind == "max_files":
        if int(payload["files_touched"]) > int(rule["limit"]):
            return rule["id"]
    elif kind == "must_include":
        if rule["token"] not in payload["stdout"]:
            return rule["id"]
    elif kind == "forbidden_path":
        for path in payload["paths"]:
            if str(path).startswith(rule["prefix"]):
                return rule["id"]
    else:
        sys.exit(f"unknown rule kind: {kind}")
    return None


def main() -> int:
    assert_frozen()
    rubric = json.loads(RUBRIC.read_text(encoding="utf-8"))
    payload = json.loads(OUTPUT.read_text(encoding="utf-8"))
    absent = missing_keys(payload)
    if absent:
        print(json.dumps({"missing": absent}))
        return 1
    failed = []
    for rule in rubric["rules"]:
        hit = check(rule, payload)
        if hit:
            failed.append(hit)
    report = {"failed": failed, "rule_count": len(rubric["rules"])}
    print(json.dumps(report, indent=2))
    return 1 if failed else 0


if __name__ == "__main__":
    raise SystemExit(main())

A matching rubric can stay this small on purpose.
Three rules are enough to prove the judge can fail.

{
  "rules": [
    {"id": "F1", "kind": "max_files", "limit": 3},
    {"id": "F2", "kind": "must_include", "token": "EXIT=0"},
    {"id": "F3", "kind": "forbidden_path", "prefix": "secrets/"}
  ]
}

Commands that enforce the order

Run the hash step before any writer process starts.
Do not skip that hash, even on a tiny fixture.
Do not let the writer process touch rubric.json at all.

mkdir -p out fixture
python3 - <<'PY'
import hashlib
import pathlib
path = pathlib.Path("rubric.json")
digest = hashlib.sha256(path.read_bytes()).hexdigest()
out = pathlib.Path("freeze.txt")
out.write_text("rubric.json %s\n" % digest, encoding="utf-8")
PY
# Writer step stays outside this page.
# Save structured writer output to out/result.json.
python3 grade.py
echo "grader_exit=$?"

That shell comment is deliberate, not an omission.
This page does not ship a live vendor client.
You already know your endpoint, or you do not.

Add a client only after the grader passes a fake file.
A fake file is a hand-built result, not a model reply.

Known pass, then known fail

Build out/result.json by hand before any model call.
You want a known pass, then a known fail.
That pair proves the judge can say both words.

{
  "files_touched": 2,
  "stdout": "built fixture\nEXIT=0\n",
  "paths": ["src/app.py", "fixture/input.txt"]
}

Change files_touched to 9 and run the grader again.
You should see rule F1 and a non-zero exit.
If that edit still exits zero, your judge is blind.

Point one path at secrets/keys and expect rule F3.
A missing fail on that path means the prefix check is dead.

Two fixtures, one judge

You do not need another model to challenge the writer.
You need two frozen inputs and one local diff.
If the outputs disagree, keep both and merge neither.

  1. Run the writer on fixture A and save the output.
  2. Run the writer on fixture B with the same rubric.
  3. Grade each output file with that same frozen rubric.
  4. Diff the two grader reports, not the two chat stories.

A disagreement is a result, not a reason to relax rules.
You inspect both diffs, then you pick neither automatically.

How you should read the exit

A clean exit only clears your frozen rule list.
It does not close design review or threat review.
You still read every changed path with your own eyes.

Result you see Move you make Reason to move
Exit 0 and the freeze still matches Keep the diff for human review The sample met rules you wrote
Exit 1 with known rule ids Repair the diff and leave the rules The writer missed a frozen check
Freeze hash does not match the file Discard the run and start again The experiment changed mid-flight
out/result.json is missing entirely Fix the writer output contract Silence must never count as a pass
Grader prints an unknown rule kind Stop and repair the rubric first A vague judge is not a judge

Refuse the self-graded demo

You should refuse any demo that grades itself.
A self-score from the writer is only a story.
Those stories can stay inside a chat window.

They are not evidence inside a pull request.
Free tokens make that story easier to produce.
That ease is the danger, not the gift itself.

Spend the free writer on alternate drafts only.
Spend your own code on the verdict every time.

A free server helps only as a process boundary.
It can run the writer away from your laptop.
It should not hold the only copy of the rubric.

If you cannot re-run the grader offline, you stop.
Offline means the rubric and the grader stay local.
A server copy alone is not a review record.

What this check will not catch

This grader checks output shape, not writer intent.
A bad rule will happily pass a bad patch.

You can freeze a mistake just as easily as a truth.
Review the rubric text before you trust any exit.

Free access can change without notice or your consent.
Do not plan a release around an unnamed quota.

Do not put production secrets inside the fixture tree.
Do not send customer rows to a trial host.

The script ignores concurrency, retries, and clock skew.
It will not catch a writer that flakes between calls.

Run the same frozen input two times in a row.
If those two outputs differ, you have no result yet.

Who should not use this

  • Skip this flow if you need a contracted uptime promise.
  • Skip it when data cannot leave your private network.
  • Skip it when nobody on the team can read the rubric.
  • Skip it when you wanted a score, not a fail list.

Also skip it for medical, legal, or payment logic.
Those domains need stronger oracles than token rules.

A string check will not save a bad payment path.
Build a real suite before you spend any model call.

One next step, then stop

Read the current MonkeyCode access terms yourself today.
Confirm the free model offer and the free server offer.
Then run this grader against a fake result file first.
Spend a writer call only after that local pass holds.

📰 Read the original article on Dev.to AI

Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.