Dev.to AI 🤖 Ai 👁 0 📖 3 min read

I Typed --help and My Agent's Script Did the Real Thing

A field note from the autonomous Claude Code agent I run every day on one Windows PC. The numbers come from its own ledgers, not from memory. This one is small, and it happened three times in 17 days. That's why it's wo

A field note from the autonomous Claude Code agent I run every day on one Windows PC. The numbers come from its own ledgers, not from memory.

This one is small, and it happened three times in 17 days. That's why it's worth writing down.

Three runs

My project has about thirty runner scripts. Most of them don't use an argument parser. They look for a handful of flags by hand, and some of them ignore everything else.

Day 1. The agent ran a mutation-testing tool with --only to check two cases. The tool had no --only. It ignored the flag and started the full list, dozens of cases, each one editing source files temporarily and running the test suite. It ran in the background. Meanwhile the agent ran other tests on top of the temporarily edited source, saw failures, and nearly blamed its own change.

Fix: that tool got argument handling. Unknown flags stop it with exit code 2 before anything runs.

Day 8. The agent wanted to see how to use another runner, so it typed --help. That runner had no --help either. It ignored the flag and started its real daily run, the one that is allowed to happen once a day. The login had expired, so no link was clicked, but with a live session it would have used up that day's run.

Fix: that runner got a list of known arguments and refuses anything else. And a written rule: in this project, read a runner's docstring instead of calling --help.

Day 17. The agent typed --help on a third runner, a reporting script. Same thing: it ran for real. This one only reads, and the working tree was unchanged afterwards, so there was no harm. It was still the same mistake.

What went wrong with the fixes

Each fix was correct and each one was local. The tool that hurt got the guard. The runner that hurt got the guard. The rule was written down, and the rule depended on remembering it.

"Ignored the flag" sounds like "nothing happened". What it really means is "something else happened": the default action, which for a runner is the real job.

The guard was added script by script, wherever a script had just caused trouble. The reporting script from day 17 still doesn't have it. Moving the guard into one shared helper that every runner calls is the obvious next step. The written rule already failed once.

The rules I'd give anyone running an agent on scripts

  • An unknown argument should stop a script before it does anything. A silent default is the worst possible response to a typo.
  • When a fix is local, list the siblings. If one script had the flaw because of how scripts in the project are written, the others have it too.
  • A lesson in a notes file is not a mechanism. The agent reads its notes, and it still repeated this one. The guard in code is the version that worked.

Where this comes from. Every post here comes from one setup I run daily: a CLAUDE.md, memory files the agent reads before it touches anything, and a separate auditor agent that returns PASS or FAIL. The first 3 chapters of the book that walks through it are free as a PDF: https://dbsoul.gumroad.com/l/autonomous-ai-agents-claude-code-free-sample

The full edition is 11 chapters plus 4 ready-to-use templates (CLAUDE.md starter, memory files, auditor checklist, measurement guide) and a hands-on section for every chapter, $19 as a PDF: https://dbsoul.gumroad.com/l/autonomous-ai-agents-claude-code

Questions about the setup are welcome in the comments — I'll answer with what actually happened, not theory.

📰 Read the original article on Dev.to AI

Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.