Rerun the Take-Home on a Clean Machine
You join the debrief with the recording still warm. The candidate walked a green script, named three files, and said the helper would flag a risky diff. You clone the repo after the call. The same command stops on a mis
You join the debrief with the recording still warm. The candidate walked a green script, named three files, and said the helper would flag a risky diff. You clone the repo after the call.
The same command stops on a missing key. The demo was a stage. The take-home was a suitcase you cannot open.
That gap is the assignment. You are not hiring a narrator. You are hiring someone who can leave a run a stranger can repeat. When the allowed path includes free model access and a free server, the stranger is you, tomorrow, on a clean machine, with none of their environment variables.
Disclosure: This article was prepared as part of MonkeyCode's product outreach. Those two availability claims, free model access and a free server, are the constraints this workflow uses.
This article does not pin model names, token quotas, hardware, duration, or permanence. Those details move. A rubric that freezes a number you have not verified will punish the wrong person. The packet should record the route the candidate actually called.
Think of the take-home as a lab sample, not a poster. A poster looks finished under good light. A sample has a label, a seal, and a way to run the same test again. You grade the seal, then the second run.
The prompt you send
You send a short assignment, not a puzzle novel. The text below is a proposal for your loop, not a transcript of a round that already happened.
Build a small review helper that reads fixtures/risky.diff and writes out/review.json. The JSON must include risk, set to low or high, and reason, one sentence. The candidate may call a model. They may run the helper on a free server if they have one.
They may not call any other network host. They commit a packet.json that tells you how to replay the run with no secrets from their laptop. Time box: four hours. You grade the replay, not a video.
That last line does the cultural work. Candidates who optimized for a screen recording will feel the floor move. You want the floor to be a command.
Own the fixture. A short diff that adds a shell call built from request input is enough. You are not asking them to invent an exploit, and you should not grade them as if they had. You are asking whether the helper flags that pattern, and whether you can reproduce the flag after they log off.
Commit this fixture yourself. Hash it before you send the repo. Your copy stays the oracle.
diff --git a/app.py b/app.py
--- a/app.py
+++ b/app.py
@@ -1,3 +1,6 @@
+import subprocess
def handle(request):
- return "ok"
+ cmd = request.args.get("cmd")
+ subprocess.run(cmd, shell=True)
+ return "ok"
The dangerous line is ordinary on purpose. Live demos love exotic bugs because exotic bugs photograph well. A replay loves a stable oracle. If the text contains a request-fed shell=True, the expected risk is high. Anything else on this fixture is a miss.
The seal on the sample
packet.json is the label. Keep it small enough to read without booking a meeting.
{
"command": "python review.py --fixture fixtures/risky.diff --out out/review.json",
"fixture_sha256": "replace-with-your-hash",
"expected_risk": "high",
"model_route": "declared-by-candidate-or-offline",
"server": "local",
"network_allow": [],
"secrets_required": false
}
secrets_required must be false. If the replay needs a key that lives only in their shell profile, the take-home is unfinished. Free model access helps only when they can point the helper at a route you are willing to supply for the replay, or when an offline rule still flags the diff.
Do not write a token budget into the rubric. Write this instead: the declared route is reachable under access you control, or the offline path passes. A number you copied from a launch post is not an oracle.
The server field is the same kind of claim. A free server is a place to run the helper, not a sticker on the README. If they used one, the packet names it and the command works there.
If they stayed local, server is local. You do not punish a local run. You punish a packet that advertises the free server while the command only succeeds on a laptop path you cannot see.
A grader you run in a sandbox
The script below is a sketch. Run it on a machine that holds no customer keys. Candidate code is still code you did not write. Path.unlink(missing_ok=True) needs Python 3.8 or newer.
#!/usr/bin/env python3
"""Proposal grader. Not a log of an executed interview."""
import hashlib, json, subprocess, sys
from pathlib import Path
def sha256(path: Path) -> str:
return hashlib.sha256(path.read_bytes()).hexdigest()
def main() -> int:
root = Path(sys.argv[1]).resolve()
packet = json.loads((root / "packet.json").read_text())
if packet.get("secrets_required"):
print("FAIL secrets_required")
return 1
fixture = root / "fixtures" / "risky.diff"
if not fixture.is_file() or sha256(fixture) != packet["fixture_sha256"]:
print("FAIL fixture hash")
return 1
out = root / "out" / "review.json"
out.unlink(missing_ok=True)
proc = subprocess.run(
["python3", "review.py", "--fixture", str(fixture), "--out", str(out)],
cwd=root,
capture_output=True,
text=True,
timeout=120,
env={"PATH": "/usr/bin:/bin", "HOME": "/tmp/empty-home"},
)
if proc.returncode != 0 or not out.is_file():
print("FAIL command", proc.returncode)
print(proc.stderr[-400:])
return 1
review = json.loads(out.read_text())
if review.get("risk") != packet["expected_risk"]:
print("FAIL risk", review.get("risk"))
return 1
reason = str(review.get("reason", "")).strip()
if not reason or len(reason) > 240:
print("FAIL reason shape")
return 1
print("PASS", packet.get("server"), packet.get("model_route"))
return 0
if __name__ == "__main__":
raise SystemExit(main())
Notice the grader does not honor packet["command"] as a shell string. You already published the command shape. Letting a submission choose a shell on your laptop is how a take-home becomes an incident. The packet command is documentation. The argument vector above is the contract.
Hash the fixture on your side, then run the grader against a checkout.
python3 - <<'PY'
import hashlib, pathlib
p = pathlib.Path("fixtures/risky.diff")
print(hashlib.sha256(p.read_bytes()).hexdigest())
PY
python3 grade_replay.py ./candidate-repo
The scrubbed environment is the analogy you should keep. You are checking the sample under your light, not under theirs. A helper that only imports because a key was exported in their terminal dies here. That death is a successful grade of the packet, even when it feels rude the first afternoon you see it.
The 120 second timeout is a gate you chose so a hung socket cannot sit in your calendar until lunch. It is not a performance score, and it is not a claim about any host. If a free server is slow, you rerun once, or you accept the offline backstop. You do not convert jitter into a character reference.
A sample solution that survives an unplugged route
A strong submission is dull. The offline rule is the backstop. A model call is optional garnish.
If you unplug the route and the file still appears with risk set to high, they understood the assignment. If the file appears only when a vendor is awake, they shipped a demo.
#!/usr/bin/env python3
"""Sample solution sketch. The rule, not the model, is the oracle."""
import json, os, sys
from pathlib import Path
def decide(diff_text: str) -> dict:
risky = (
"subprocess" in diff_text
and "shell=True" in diff_text
and "request" in diff_text
)
if risky:
return {
"risk": "high",
"reason": "shell command built from request input",
}
return {
"risk": "low",
"reason": "no request-fed shell call in the fixture",
}
def main() -> int:
args = sys.argv
fixture = Path(args[args.index("--fixture") + 1])
out = Path(args[args.index("--out") + 1])
review = decide(fixture.read_text())
route = os.environ.get("MODEL_ROUTE", "").strip()
if route:
review["model_route_declared"] = route
out.parent.mkdir(parents=True, exist_ok=True)
out.write_text(json.dumps(review))
return 0
if __name__ == "__main__":
raise SystemExit(main())
Tell them a model may disagree with the rule, and the rule wins on this fixture. That sentence saves you from a confident paragraph that marks the diff low because a nearby comment mentioned tests. Fluency is cheap. The oracle is not.
If they declare a model route, they also declare it was not required for the pass. That is the grown-up use of free model access. The free path can draft the reason sentence. It cannot be the only reason the JSON exists.
A free server, when they used one, gets the same treatment. Name it, prove the command there, or write local and move on. A hostname with no replay behind it is a slide.
This heuristic is a toy oracle for one file. Do not describe it as a scanner, and do not hire someone because three in checks impressed you. You are grading whether they can separate a stable check from a live demo. A real review bot needs a wider net than this fixture. Say that in the invite so a clever candidate does not waste the four hours building a framework you will not run.
How you say the score out loud
You do not need a spreadsheet if you say the same four checks every time.
The seal comes first. The fixture hash matches, secrets are not required, and the command shape is the one you published. A packet that says "see the video" scores zero on this check. You already saw the video. It did not clone.
The second run comes next. Exit code zero, out/review.json appears, and risk matches your oracle. You do not grade the reason for wit. You grade that it is one sentence and that it names the shell pattern. A short novel in that field is a smell, not a bonus.
The boundary comes third. model_route is whatever they called, including offline. network_allow stays empty unless you agreed to a host. A free server earns credit only when server matches where the command ran. A mismatch is a fail. A local run is a pass on this check.
Wait, that last pair is five beats. Say it as two breaths instead. Match the field to the place the command ran. Local is allowed. A lie about the host is not.
Restraint comes last. No extra hosts, no socket opened just to log, no pastebin call hiding in review.py. The helper that phones home is a different product from the helper you asked for. You are allowed to stop reading there.
After a pass, keep three artifacts: the packet, the fixture hash, and the review.json your grader wrote. Drop the screen recording as evidence. If a committee asks how you know, you show the command and the PASS line. That argument is smaller than a fourteen-minute video, and you can repeat it when someone challenges the score six weeks later.
Failure modes that show up after the smile
The haunted laptop is the common one. The README is confident. The client builds itself at import time and demands a key.
Your scrubbed environment never reaches decide. Treat import-time network as a fail. A module that cannot be imported without a secret cannot be replayed, no matter how smooth the call sounded.
The trophy server is the next one. The packet names a free host. Nothing in the repo contacts that host, and your argument-vector run is purely local.
They used the server as a slide. If your replay will not hit the host, the honest field is local. Slides do not get points.
The moving fixture shows up when a weak heuristic cannot see shell=True, so they edit your diff and rehash it. Your copy is the source of truth. Hash mismatch is not a conversation. It is a different assignment, and you did not ask for a different assignment.
The model monologue is subtler. The second run exits zero. risk is low. reason is a long page of generic advice about pinning dependencies.
The command worked, and the judgment is wrong. Your expected risk outvotes the prose. Video grading cannot see this failure, because the prose sounds like seniority.
The last failure is yours. You encode a quota, a brand, or a latency cap you have not verified, then fail someone because a free path was slow on a Tuesday. Free access is a constraint you offer, not a benchmark you publish. If the route is down, rerun later or accept the offline backstop. Do not write the outage into their score.
Who should skip this gate
Skip it when the role is a design hire and the work product is a diagram. A replay will not show taste, and pretending it will waste a week. Skip it when you cannot run stranger code in a sandbox. The grader is a footgun on a machine that holds production tokens, even with the argument vector locked.
Skip it when your loop is a forty-minute pairing session and you will never clone the repo. Sending this prompt and then grading the recording teaches people that the packet was theater.
Also skip it if you wanted a leaderboard. One fixture, one expected risk, one replay: that is a gate, not a ranking.
Two candidates who pass are not ordered by this script. They both left a sample you could open. Order them with a different exercise, on a different day, about a different skill.
The line you leave in the invite
Close the invite with one practical sentence, not a pitch. If a free model path and a free server are on offer through MonkeyCode, the candidate may use either, and packet.json has to show which path the replay used. The offer is a constraint, the way a library card is a constraint. It does not grade itself, and it does not expire on a timer you invented in the rubric.
You still read the reason sentence. You still care that they noticed the shell call and not a typo in a comment. The replay only stops you from hiring the recording. The recording was never the colleague who has to debug this helper on a Monday when the route is down.
Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.