Dev.to AI 🤖 Ai 👁 0 📖 9 min read

Leave a Trace a Classmate Can Run: A Bootcamp Lab on Checkpoint Replay

If your checkpoint log will not replay, the lab is not done. A polished portfolio screenshot will not talk me out of that. People are busy arguing whether a vibe-coded site still counts as a portfolio. Fair question. Wr

If your checkpoint log will not replay, the lab is not done. A polished portfolio screenshot will not talk me out of that.

People are busy arguing whether a vibe-coded site still counts as a portfolio. Fair question. Wrong assignment for this week. I care whether the next student can rerun your setup without borrowing your laptop.

The grade is a replay, not a vibe

I do not grade the homepage. I grade a trace.

You ship a bootstrap, one JSONL log, and a checker I can run with the network off. That is the whole submission. The page, if you built one, is an exhibit. Exhibits do not pass labs.

Why so strict? Bootcamp demos fail in a boring, expensive way. The terminal looked finished at 4pm, the box was reset by morning, and the draft that "worked" was never saved. I do not want that failure mode in the gradebook.

So the lab asks a rude question. Can I clone your repo and prove the setup happened, in order, without you in the room?

What you set up before any model call

Keep the tree dull. Dull is gradable.

lab/
  fixtures/tiny.md
  expected/checkpoints.json
  scripts/bootstrap.sh
  scripts/replay-check.mjs
  runs/.gitkeep
  README.md

fixtures/tiny.md is about twenty lines. A heading, a bug note, one expected phrase. Not a startup. If the fixture is a novel, you are hiding from the assignment.

The contract is a list, not a mood:

{
  "required": [
    "prereq_ok",
    "fixture_loaded",
    "draft_requested",
    "oracle_compared",
    "log_sealed"
  ]
}

Order is the point. Sealing the log before the fixture loads is fiction. A later reader should not have to trust your memory of which step came first.

Bootstrap stays boring on purpose.

#!/usr/bin/env bash
set -euo pipefail
cd "$(dirname "$0")/.."
test -f fixtures/tiny.md
mkdir -p runs
node --version >/dev/null
echo "prereq_ok"

Run it. Run it again. If the second run needs you to delete a folder by hand, you do not have a setup. You have a ritual.

I also want a missing-file drill before anyone celebrates. Rename the fixture, run bootstrap, put the file back. Did the script fail before any draft step? If it still "succeeds," the check is optimistic, and optimism is not a checkpoint.

Two roles, and they do not share a fate

Students may draft the script with a model. They may also run the lab on a remote box, so the grade does not depend on a personal laptop nobody else can reset.

Disclosure: This article was prepared as part of MonkeyCode's product outreach.

The only product facts I am using are these: MonkeyCode offers free model access, and it offers a free server option. I am not stating a token quota, a hardware profile, a session length, a model name, or a promise that free access remains. Those details move. If a syllabus needs them, check the current docs and write down only what you verified.

Here is why those two options belong in this lab at all.

  • A free model may draft bootstrap.sh. It may not grade the lab.
  • A free server may be the student box you reset between attempts.
  • The grade runs on replay-check.mjs against the committed log. Offline. Local. No model in the loop.

What if free access is down during class? You write the script yourself, set "assisted": false, and keep going. The checker never phones home. Assistance is a convenience. The trace is the assignment.

That split is the lesson. Not a slogan. A boundary. Who typed the script can change. The order of evidence cannot.

The checker you will actually submit

Treat the following as a proposed scaffold, not as a timed benchmark. I am not publishing latency numbers. A green run is not a model evaluation, and it should not be quoted like one.

import { readFileSync } from "node:fs";

const expected = JSON.parse(
  readFileSync(new URL("../expected/checkpoints.json", import.meta.url))
);
const logPath = process.argv[2];
if (!logPath) {
  console.error("usage: node scripts/replay-check.mjs runs/latest.jsonl");
  process.exit(2);
}

const events = readFileSync(logPath, "utf8")
  .trim()
  .split("\n")
  .filter(Boolean)
  .map((line, index) => {
    const row = JSON.parse(line);
    if (!row.name || !row.at) {
      throw new Error(`line ${index + 1} missing name or at`);
    }
    return row;
  });

const names = events.map((event) => event.name);
let cursor = 0;
for (const required of expected.required) {
  const found = names.indexOf(required, cursor);
  if (found === -1) {
    console.error(`missing checkpoint: ${required}`);
    process.exit(1);
  }
  cursor = found + 1;
}

if (events.at(-1).name !== "log_sealed") {
  console.error("log is not sealed");
  process.exit(1);
}
console.log(`replay_ok events=${events.length}`);

A passing log is plain. Timestamps below are examples for the fixture, not a claim that I ran this on a particular day in production.

{"name":"prereq_ok","at":"2026-10-10T15:00:00Z"}
{"name":"fixture_loaded","at":"2026-10-10T15:00:01Z","bytes":412}
{"name":"draft_requested","at":"2026-10-10T15:00:04Z","assisted":true}
{"name":"oracle_compared","at":"2026-10-10T15:00:05Z","pass":true}
{"name":"log_sealed","at":"2026-10-10T15:00:06Z"}

Notice what I refused to store. No model essay. No API key. No "the UI felt modern." The local oracle is a student-written check against tiny.md, recorded as pass. If the model grades its own draft, you have wandered into a different course. Come back.

Appending a line can be ugly. Ugly is readable.

now=$(date -u +%Y-%m-%dT%H:%M:%SZ)
printf '%s\n' "{\"name\":\"fixture_loaded\",\"at\":\"$now\",\"bytes\":412}" >> runs/latest.jsonl

On macOS, date -u is enough for this format. If the image is stripped down, write the timestamp from Node instead of fighting the shell. Same log either way. Do not invent a second schema just to feel modern.

Commands the README must show

I do not accept "run the project" as a step. I accept copy-paste. A classmate should not have to guess your directory layout.

git clone <your-repo> lab-replay && cd lab-replay
bash scripts/bootstrap.sh
node scripts/replay-check.mjs runs/latest.jsonl

Expected last line: replay_ok events=5. If your count is higher because you logged extra notes, say so in the README. Extra events are fine. Missing required ones are not.

What if node is missing on the free server image? Bootstrap should fail at the prerequisite check, not halfway through a draft. Write that failure in one README line. Classmates hit missing runtimes. Pretending they will not is how labs melt.

Stops I walk in class

Four stops. Each one has a question you should answer without slides.

  1. Clean tree. Does bootstrap exit 0 on a fresh clone, then again with no cleanup? If not, why are you already drafting prompts?
  2. Fixture. Does the log store a byte count? If the file is missing, does the script die before draft_requested?
  3. Help flag. Did a model touch the script? Then assisted is true. Did you type it? Then false. Either can pass. A missing flag cannot.
  4. Airplane mode. I copy the repo and run the checker offline. If it tries to call a model, grading stops. Not "partial." Stops.

At stop 3 I ask one spoken sentence. What did you keep from the draft, and what did you throw away? The script cannot hear you. I can. A green log plus a blank stare becomes a review note, not an automatic pass, and not a silent zero either.

A failure I want practiced, not theorized

Rename fixtures/tiny.md to fixtures/tiny.md.bak. Run bootstrap. You should stop. Write nothing that claims fixture_loaded. Restore the file. Append the real events only after the file is back.

Then aim the checker at a bad log on purpose:

node scripts/replay-check.mjs runs/failure.jsonl
echo "exit=$?"

I want missing checkpoint: fixture_loaded and a non-zero exit. If your checker prints a stack trace and still exits 0, the catch path is theater. Theater does not replay.

Questions I will ask while you do this: where did the script notice the missing file? Did the model draft hide that check inside a comment? Are you logging the absence, or only the happy path?

Happy-path-only logs are how screenshots happen. I am allergic to them this week.

Stretch, after the core is green

Do not start here. I mean it.

  • Commit runs/failure.jsonl where fixture_loaded never appears. The checker must exit 1 and name the gap.
  • Hand the repo to a classmate. README only. Ten minutes. No chat.
  • Produce two logs from the same contract: one assisted, one hand-written. The checker should not care who typed the script.
  • If you also shipped a portfolio page, link it under "exhibit." It adds zero points unless replay still passes.

That last stretch is my small answer to the portfolio argument. A generated site can be worth showing. It is not evidence that your setup exists.

Rubric I can defend to the class

Slow model calls are not a character flaw. Fake flags are.

Row Points What pass means
Idempotent bootstrap 25 Second run exits 0 with no hand cleanup
Ordered checkpoints 25 Checker prints replay_ok
Honest failure path 20 Missing fixture exits 1 and names the event
Honest assisted flag 15 Flag matches the work you actually did
README a classmate can follow 15 They finish without pinging you

The order row is all or nothing. You do not get partial credit for three of five events. Why? A skipped fixture with a sealed log is exactly the demo I am refusing to trust.

The README row allows partial credit. Seven minutes and one stuck command can be a 10, not a 0, if the classmate still finishes from your notes.

A late free-server reset does not deduct points if the committed log still replays on my machine. You lost a box. You did not lose the trace. That is the fair part. Availability wobble is not the same thing as a missing fixture.

Who should skip this lab

This will not make you hireable by itself. It will not audit CSS, and it will not tell you which model is best. I have not measured that, so I will not fake a ranking.

Skip it when:

  • The week is visual design, accessibility, or a public portfolio critique
  • You need a named model, a quota, a GPU shape, or a permanence clause in the syllabus
  • The log would contain secrets, tokens, or another student's private notes
  • You wanted a leaderboard — this checker will not rank models, and I will not add that

A green replay proves one claim: this repo can retell its setup in order. It does not prove you understand the code a model drafted. That is a real hole. The spoken sentence at stop 3 is how I patch it, and it is a human patch. Do not automate it into a fake sentiment score.

I also would not drop this rubric onto a hiring review. Hiring is a different promise. This promise is narrower: leave a trace the next person can run.

One way to try the boundary

If you run a lab this week, put drafts and remote boxes on the student side. Keep the grade on the log. MonkeyCode's free model access and free server option are a practical way to try that split, provided you treat them as availability, not as the rubric.

📰 Read the original article on Dev.to AI

Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.