A company of AI agents that actually finishes things: board, handoffs, a gate, and a meter
I kept rebuilding the same failure: five agents in one chat window, no memory of what was blocked, and a "done" that nobody had checked. Better prompts did not fix it. A bigger model did not fix it either. Structure did
I kept rebuilding the same failure: five agents in one chat window, no memory of what was blocked, and a "done" that nobody had checked.
Better prompts did not fix it. A bigger model did not fix it either. Structure did. A company has a board, roles with defined handoffs, a gate before anything ships, and a budget someone is actually watching. So I wrote that scaffolding down as code instead of vibes: four Python modules, standard library only, plus a five-role roster and a demo you can run offline, with no API keys, in seconds.
The failure mode, stated precisely
Four symptoms, all of them organizational rather than model-related:
- Nothing remembers what is blocked. Agent 3 asks for something Agent 1 already produced, or waits forever on a step that silently failed.
- "Done" means "the model stopped talking." No check between an artifact being produced and an artifact being shipped.
- Loops. A task blocks, gets retried with the same input, blocks again, forever. Costly and invisible.
- Nobody watches the meter. You find out you burned the month's quota on the 31st, when it is a post-mortem instead of a control.
Each of those is a missing structure, not a missing prompt. Here is what I put in its place.
1. The board — tasks with real dependencies
A SQLite table, a parents list per task, and one transition rule: a task becomes ready only when every parent is done. No prose ordering, no "we'll do this after that, remember".
def refresh(self):
"""todo -> ready once every parent is done; done parents unlock children."""
rows = self.conn.execute("SELECT id,status,parents FROM tasks").fetchall()
done = {r["id"] for r in rows if r["status"] == "done"}
for r in rows:
if r["status"] != "todo":
continue
parents = json.loads(r["parents"])
if all(p in done for p in parents):
self.conn.execute(
"UPDATE tasks SET status='ready',updated_at=? WHERE id=?", (_now(), r["id"])
)
self._event(r["id"], "ready", {"unlocked_by": parents})
Every transition — created, ready, claimed, completed, blocked — appends a row to an events table. The log is append-only, so a crash never erases history, and "why is this task not moving?" is one query instead of a slack thread.
Statuses are a small, boring set: todo | ready | running | blocked | review | done. That is the whole state machine. It fits on a sticky note, which is the point — an agent team does not need a workflow engine, it needs an unambiguous answer to "can this start yet?".
2. The dispatcher — routing order, and the anti-loop rule
The dispatcher is a dict of {assignee: handler} plus a routing order (the org chart, in code). Each pass it asks the board for ready tasks per assignee, claims them, runs the handler, and records the handoff.
Handlers are just callables that return a dict:
handlers = {
"scout": with_board(scout),
"analyst": with_board(analyst),
"builder": with_board(builder),
"auditor": with_board(auditor_handler),
}
disp = Dispatcher(board, handlers, order=("scout", "analyst", "builder", "auditor"))
trace = disp.run(max_cycles=10)
Two rules that did most of the work for me:
- A handler that raises blocks the task (with the error in the event log) instead of killing the run. One broken stage does not take the pipeline down.
-
A repeated block escalates instead of retrying. Retry policies hide broken work; escalation surfaces it. The dispatcher carries a
block_limit; the policy we actually run is: after N blocks on the same task, stop and route it to a human. Note the honest version of that seam — the shipped offline demo drives the happy path and does not push a task into repeated blocks, so the escalation counter is the one piece the demo leaves for you to wire to your own paging/alerting.
3. The gate — two layers, deterministic
This is the part that changes the outcome. A contract check that is structural, and an audit check that is substantive.
Layer one: the handoff contract. Per role, the keys that must be present on completion. Not "did it produce text", but "did the analyst actually emit a shortlist and a recommendation".
HANDOFF_CONTRACT = {
"scout": ["candidates", "sources", "confidence"],
"analyst": ["shortlist", "recommendation", "sources", "confidence"],
"builder": ["offer", "sources", "confidence"],
"marketer": ["channels", "drafts"],
"auditor": ["verdict", "issues", "checked"],
}
An empty list is a legitimate value (an auditor with no issues); only a genuinely absent or blank key counts as missing. That distinction matters more than it looks: a gate that flags "no issues found" as a failure trains everyone to ignore the gate.
Layer two: the audit gate. It re-derives the arithmetic, requires a URL and a date on every cited source, and rejects numbers that are not labelled. It is deterministic — no LLM call — so a run is reproducible and the verdict is not a coin flip.
if None not in (price, volume, gross):
calc = round(float(price) * float(volume), 2)
if abs(calc - float(gross)) > 0.005:
issues.append({"severity": "high",
"what": "gross mismatch: %s x %s = %s but artifact says %s"
% (price, volume, calc, gross)})
...
verdict = "changes" if any(i["severity"] in ("high", "med") for i in issues) else "approve"
The URL-and-date requirement is the cheapest honesty mechanism I have found. 29 EUR (checked 2026-10-06, <link>) survives a review; ~29 EUR quietly becomes a fact three documents later, with no one able to say where it came from.
4. The meter — the budget rule on the critical path
Every agent step appends one JSON line to usage.log (append-only; a crash never loses history). The report answers one question: is spend inside the month's pace, or is the month being burned in the first week?
def report(self, days_elapsed):
"""Pace rule: spend% must stay <= proportional target (1 - days_left/month)."""
spent = self.total()
target_pct = 100.0 * (days_elapsed / self.days_in_month)
spent_pct = 100.0 * (spent / self.monthly_usd) if self.monthly_usd else 0.0
status = "OK"
if spent_pct > target_pct + 10:
status = "AHEAD"
elif spent_pct > target_pct:
status = "WATCH"
return {"spent_usd": spent, "monthly_usd": self.monthly_usd,
"spent_pct": round(spent_pct, 1), "target_pct": round(target_pct, 1),
"status": status}
OK / WATCH / AHEAD is deliberately coarse. A precise number nobody reads loses to three words that show up in a log every day.
The part most people skip: the negative control
A gate that has never failed proves nothing. So the demo feeds the audit gate a deliberately corrupted artifact — one where the stated gross does not equal price x volume — and checks that the gate says so.
Here is the actual output of python3 demo/run_demo.py (task ids are uuids regenerated each run; the rest is verbatim):
Board seeded: 4 tasks, 3 dependency edges.
goal <id>
analyst <id> (needs goal)
builder <id> (needs analyst)
auditor <id> (needs builder)
Dispatcher trace (4 transitions):
<id> -> done
<id> -> done
<id> -> done
<id> -> done
Handoff contract check:
scout OK
analyst OK
builder OK
auditor OK
Audit gate verdict: approve
checked: price x volume arithmetic, source url+date, claim labels, builder handoff contract: ok
Usage meter (day 3 of 31):
{'spent_usd': 1.16, 'monthly_usd': 100.0, 'spent_pct': 1.2, 'target_pct': 9.7, 'status': 'OK'}
Gate negative control (deliberate gross error): changes -> gross mismatch: 49.0 x 42 = 2058.0 but artifact says 999.0
DEMO RESULT: PASS - full pipeline ran offline (+ gate catches errors)
The meter line is demo data from four mock steps (0.12 + 0.35 + 0.48 + 0.21 = 1.16 USD), not a benchmark of anything. The interesting line is the second-to-last one: approving good work is easy; the value of the gate is that it refuses.
Three things I'd tell anyone building this
- Escalate, don't retry. Retry policies hide broken work; escalation surfaces it. Cap failed blocks and route them to a human, then fix the task, not the loop.
- Label every number. Source, date, or the word "estimate". The alternative is a number that outlives its evidence.
- Put the meter on the critical path. A budget report you read on the last day of the month is a post-mortem, not a control.
What this is not
It is a skeleton, not a platform. The demo runs deterministic mock handlers so it is reproducible — swapping them for calls to any OpenAI-compatible endpoint does not change the board, the gate or the meter. There is no server to run, no dashboard, and no lock-in: it is source you own and edit. Python 3.8+, standard library only. The demo needs no API key and no account.
The five role definitions (orchestrator, scout, analyst, builder, auditor) ship as markdown with frontmatter, so the "org chart" is a file you can argue with in a PR.
Disclosure: I packaged this skeleton into a kit and I sell it — the paid version is my own product, which is how the time I spend on this gets funded. Everything described in this post is the engine: board, handoffs, gate, meter, demo. The kit adds the written operating playbook (pipeline playbooks, research standards, a compliance checklist), three extra pipeline packs, a standalone quota-pace script and lifetime updates: Agent Company Kit on Gumroad. If you'd rather build the four parts yourself from the code above — that is a perfectly good outcome, and honestly the reason I wrote this post.
Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.