Workshop: Classify Endpoint Failures Before a Second Call in 70 Minutes
A second model call is a budget decision, not a reflex, and this seventy-minute workshop teaches a local classifier that blocks the reflex. Students replay seven recorded failures, assign retry once, stop, or inspect, an
A second model call is a budget decision, not a reflex, and this seventy-minute workshop teaches a local classifier that blocks the reflex. Students replay seven recorded failures, assign retry once, stop, or inspect, and leave with a standard-library script that runs without a network. The exercise fits a shared lab that already calls a free model endpoint and may attach a free server for class time. It does not prove that any host will keep those options, publish quotas, or name a stable model.
Why the label comes before the dashboard
The useful question is narrower than a general retry slogan, and it shows up whenever a lab client treats every error as temporary. A free endpoint still spends shared capacity and attention when the client repeats a request that the contract already rejected. A transport timeout and a schema rejection both print as failures, yet only the timeout can justify one more attempt. The workshop therefore grades the failure label, not the prose quality of whatever answer a model might have returned.
Outcomes students must be able to defend
By minute seventy, each student should be able to defend three actions with a record, a reason, and a second-call boolean. Retry once means a single extra attempt is allowed, and the next identical failure must stop. Stop means the client must not call again until a human changes the request, the credential, or the contract. Inspect means the row is incomplete, so automation should surface it instead of guessing a retry.
Clock for a prepared room
The clock assumes a prepared room, so instructors should paste the script into the workspace before students arrive. Each block has an exit check, and the class should not advance while that check is still red. Ten minutes of slack sits inside the middle blocks, which is enough to debug a path error or a Python version mismatch. The final ten minutes are for limitations, not for a live demo that the room has not measured.
- Minutes 0-10: define retry once, stop, and inspect without opening the script.
- Minutes 10-25: run the module and match the designed seven-line summary.
- Minutes 25-45: write a defense for the schema row and the rate-limit row.
- Minutes 45-60: trace the dry-run guard and add one unclassified status.
- Minutes 60-70: submit a memo listing three cases the script must not decide.
Materials and honesty rules
The graded path uses Python 3.11 or newer and no third-party packages, which keeps the room independent of package mirrors. Students need a terminal, a text editor, and the fixture list inside the module they are about to save. A live key is out of scope, because the point is to classify records the instructor already captured or composed. Treat the script as a teaching artifact with a designed exit check, not as a measured benchmark from a production host.
Decision table students implement
The matrix below is the contract students implement, and the script is only a runnable copy of that contract. Status numbers are a teaching subset, so an unlisted code must land in inspect rather than in a silent default retry. A 429 is inspect on purpose, because a free shared pool may be enforcing a limit rather than reporting a short blip. Connection failures stay eligible for one retry, and the already-retried flag then forces stop on the following row.
| Signal | Action | Second call | Note |
|---|---|---|---|
| error is timeout, connection_reset, or dns_temporary | retry_once | yes, once | Transport may clear without a contract edit |
| status 408, 500, 502, 503, or 504 | retry_once | yes, once | Stop instead if already_retried is true |
| already_retried is true | stop | no | The one-attempt teaching budget is spent |
| status 400, 401, 403, 404, 413, or 422 | stop | no | The same call will not repair the contract |
| status 429 | inspect | no | Do not inherit the transient retry path |
| status 200 with an empty body | inspect | no | A success status is not a complete answer |
| missing or unlisted status | inspect | no | Guessing a retry is a failed exit check |
- The classifier checks already_retried before status, so a repeated 503 stops because the one-attempt budget is already spent.
- The classifier checks coarse transport errors before status, so a timeout with a missing status still qualifies for one retry.
- The classifier checks 429 before the transient set, so a rate-limit row cannot slip into an automatic second call.
Worked module students rerun
Save the following module as retry_gate.py and run it with Python 3 before the discussion period begins. The example is local and unexecuted against any vendor, and the printed lines are the designed checks students should reproduce. If a row fails, fix the classifier rather than editing the expected label to match a convenient result. That rule keeps the lab honest when a later cohort wants to add records from a real log.
#!/usr/bin/env python3
'''Teaching fixture: classify a failure before a second model call.
Local and unexecuted against any vendor. Designed checks are not a benchmark.
'''
from __future__ import annotations
import sys
from dataclasses import dataclass
RETRY_ONCE = 'retry_once'
STOP = 'stop'
INSPECT = 'inspect'
STOP_STATUS = {400, 401, 403, 404, 413, 422}
RETRY_STATUS = {408, 500, 502, 503, 504}
TRANSPORT_ERRORS = {'timeout', 'connection_reset', 'dns_temporary'}
@dataclass(frozen=True)
class Decision:
action: str
reason: str
def classify(record: dict) -> Decision:
status = record.get('status')
error = (record.get('error') or '').lower()
body = record.get('body')
if record.get('already_retried'):
return Decision(STOP, 'retry budget already spent')
if error in TRANSPORT_ERRORS:
return Decision(RETRY_ONCE, 'transport error: %s' % error)
if status in STOP_STATUS:
return Decision(STOP, 'client or auth status %s' % status)
if status == 429:
return Decision(INSPECT, 'rate limit needs a human rule')
if status in RETRY_STATUS:
return Decision(RETRY_ONCE, 'transient status %s' % status)
if status == 200 and not body:
return Decision(INSPECT, 'empty success body')
if status is None:
return Decision(INSPECT, 'missing status')
return Decision(INSPECT, 'unclassified status %s' % status)
def allow_second_call(record: dict) -> bool:
decision = classify(record)
return decision.action == RETRY_ONCE and not record.get('already_retried')
def plan_next(record: dict) -> dict:
decision = classify(record)
if allow_second_call(record):
return {
'id': record['id'],
'plan': 'one_retry',
'reason': decision.reason,
}
return {
'id': record['id'],
'plan': 'skip',
'reason': decision.reason,
}
FIXTURES = [
{
'id': 't1',
'status': 503,
'error': '',
'body': None,
'already_retried': False,
'expect': RETRY_ONCE,
'expect_plan': 'one_retry',
},
{
'id': 't2',
'status': 401,
'error': '',
'body': {'message': 'unauthorized'},
'already_retried': False,
'expect': STOP,
'expect_plan': 'skip',
},
{
'id': 't3',
'status': None,
'error': 'timeout',
'body': None,
'already_retried': False,
'expect': RETRY_ONCE,
'expect_plan': 'one_retry',
},
{
'id': 't4',
'status': 429,
'error': '',
'body': None,
'already_retried': False,
'expect': INSPECT,
'expect_plan': 'skip',
},
{
'id': 't5',
'status': 503,
'error': '',
'body': None,
'already_retried': True,
'expect': STOP,
'expect_plan': 'skip',
},
{
'id': 't6',
'status': 200,
'error': '',
'body': None,
'already_retried': False,
'expect': INSPECT,
'expect_plan': 'skip',
},
{
'id': 't7',
'status': 422,
'error': '',
'body': {'message': 'schema'},
'already_retried': False,
'expect': STOP,
'expect_plan': 'skip',
},
]
def run() -> int:
failures = 0
for row in FIXTURES:
decision = classify(row)
plan = plan_next(row)
action_ok = decision.action == row['expect']
plan_ok = plan['plan'] == row['expect_plan']
if not (action_ok and plan_ok):
failures += 1
print(
'mismatch id=%s expected_action=%s expected_plan=%s'
% (row['id'], row['expect'], row['expect_plan']),
file=sys.stderr,
)
second = allow_second_call(row)
print(
'%s: action=%s second_call=%s plan=%s reason=%s'
% (row['id'], decision.action, second, plan['plan'], decision.reason)
)
print('checked=%s failed=%s' % (len(FIXTURES), failures))
return 1 if failures else 0
if __name__ == '__main__':
sys.exit(run())
python3 retry_gate.py
echo exit:$?
How to read the designed output
Read the designed output as a contract for the checker, not as a log captured from a running server. True on second_call appears only for t1 and t3, and every other row must print False and plan skip. Row t5 is the budget lesson, because the status is still 503 while already_retried forces an immediate stop. If your local run differs, treat the difference as a bug in the copy, not as a new research finding about endpoints.
t1: action=retry_once second_call=True plan=one_retry reason=transient status 503
t2: action=stop second_call=False plan=skip reason=client or auth status 401
t3: action=retry_once second_call=True plan=one_retry reason=transport error: timeout
t4: action=inspect second_call=False plan=skip reason=rate limit needs a human rule
t5: action=stop second_call=False plan=skip reason=retry budget already spent
t6: action=inspect second_call=False plan=skip reason=empty success body
t7: action=stop second_call=False plan=skip reason=client or auth status 422
checked=7 failed=0
The designed process exit is 0. That number only means the local expectations matched the classifier, and it says nothing about latency, token use, or host capacity.
Guided exercises
Minutes 0 to 10: name the actions
Ask each pair to rewrite the three definitions in their own words without looking at the script. Collect one example of a stop case from ordinary HTTP practice, such as a rejected schema or a missing credential. Do not accept a complaint about model quality as a class, because that sentence does not tell the client what to do next. End this block only when every pair can point to the boolean that would allow a second call.
Minutes 10 to 25: run the deck
Students run the module and compare seven lines with the decision table, including the second-call column. The designed summary is checked equals 7 and failed equals 0, which is a specification rather than a published performance number. If the room has no Python, an instructor can still walk the table row by row and mark the same booleans on paper. That paper path is slower, but it preserves the lesson when install rights are locked down by the lab image.
Minutes 25 to 45: defend stop and inspect
Each student picks the 422 row and writes why a second identical call cannot repair a schema mismatch. Each student also picks the 429 row and writes why an automatic retry could amplify load on a shared free pool. The discussion should stay on client behavior, not on blaming a model for a status code the client already understands. A useful note names the field that carried the evidence, such as status, error, or already_retried.
Minutes 45 to 60: trace the guard and add t8
Students trace plan_next and confirm that skip is the plan for every row except the first 503 and the timeout. They then add a t8 fixture with status 409, an expect action of inspect, and an expect plan of skip. The existing fall-through rule should accept that row without a new branch, which is the point of the unclassified path.
After that row is appended correctly, the designed summary becomes checked equals 8 and failed equals 0. No socket opens in this block, so a mistaken label cannot trigger a live call while the class is still learning. Keep the new fixture in the same file so the next cohort can rerun the extended deck without a separate handout.
{
'id': 't8',
'status': 409,
'error': '',
'body': {'message': 'conflict'},
'already_retried': False,
'expect': 'inspect',
'expect_plan': 'skip',
}
Minutes 60 to 70: write the limitation memo
Each student writes three cases the script must not decide, using complete sentences and naming the missing field. Acceptable cases include a non-idempotent tool call, a 429 that includes Retry-After, and a vendor code outside the teaching set. Unacceptable cases are vague complaints about model quality, cost, or speed, because this checker never saw those signals. Collect the memos before anyone opens a host dashboard, so the product discussion cannot replace the exit check.
Map one synthetic line
Use a synthetic line, not a production export, when you demonstrate the mapping on a projector. A synthetic line with id lab-17, status 503, an empty error, and retried no becomes the same shape as fixture t1. Students can append that mapped object to FIXTURES only after they also write the expected action and expected plan. If they forget the expected fields, the checker should fail closed rather than treat a missing expectation as success.
Where a host with free access fits
MonkeyCode enters the lesson only after the labels are stable, and only as one possible lab host. Disclosure: This article was prepared as part of MonkeyCode's product outreach. This article relies on two availability claims supplied for the draft: free model access and a free server option. This draft does not state quotas, hardware, duration, uptime, or benchmark scores, because those facts were not supplied for verification.
Students should read the host's current documentation before class, since availability claims can change and this page is not that documentation. A free server option may help a classroom share one practice host, but it does not change the failure labels or the one-retry budget. If documentation contradicts this workshop, follow the documentation and keep the local fixtures as the graded artifact.
Export rules if logs come later
If the instructor later exports logs, keep a local id, status, a coarse error tag, and an empty-body flag. Drop raw prompts, completions, authorization headers, and user identifiers before the file reaches any student machine. Map vendor error strings to a small tag set such as timeout, connection_reset, or dns_temporary, and leave unknown strings empty. An unknown string must not become a retry, because the classifier treats an empty error plus an unlisted status as inspect.
Limitations of the teaching contract
The status sets are incomplete, and a 409, 425, or vendor-specific code will fall through to inspect until an instructor adds a reviewed rule. One retry can still be the wrong policy for 503 storms, so the flag is a teaching brake rather than a production backoff schedule. The script does not measure tokens, latency, or queue depth, and it should not be cited as evidence about cost or speed. Empty body detection is naive, because a JSON null, a whitespace string, and a missing key are different defects in a real client.
A 429 with a Retry-After hint might be safe to defer, but these fixtures carry no such hint, so the honest label is inspect. Idempotency is assumed for the one retry, and a non-idempotent tool call should be stop even when the status looks transient. That case is listed in the closing memo because the script cannot see side effects that were not recorded in the row.
Who should not use this lab
Skip the lab if the course outcome is a tuned retry library with jitter, circuit breakers, and idempotency keys. Skip it if students must handle protected data, because the safe path here assumes fixtures that contain no personal content. Skip it if the goal is to compare model quality, since none of the rows include a completion to score. On-call playbooks also need vendor documents and ownership, which this seventy-minute room deliberately does not provide.
Facilitation notes
Instructors should resist turning the closing memo into a tool bake-off, because the room has no measured vendor results to compare. Grade the written reasons, the passing local checker, and the limitation list, not the number of live calls a student managed to send. Pairs are enough; a larger group hides whether each person can name the field that justified stop. If Python is unavailable, the paper walk still counts, provided the booleans match the table exactly.
Closing check
The core result is testable: seven local rows, three actions, and a guard that refuses an unearned second call. Run the module, confirm the designed summary, and only then consider pointing the same field map at a host log. If a lab already uses MonkeyCode's free model access and free server option, classify redacted error rows with these labels first. That pass is a review step, not a claim that the host produced the sample rows in this article.
Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.