My agent guardrail was fine. Its kubectl parser wasn't.
I maintain Aegis-DevOps, an open-source check that sits between an AI coding agent and the shell. Before Claude Code, Codex, Copilot, Cursor or Gemini CLI runs kubectl, terraform or a cloud CLI, Aegis turns the command i
I maintain Aegis-DevOps, an open-source check that sits between an AI coding agent and the shell. Before Claude Code, Codex, Copilot, Cursor or Gemini CLI runs kubectl, terraform or a cloud CLI, Aegis turns the command into an intent (action, resource, namespace) and matches it against signed policy. Recently I asked three AI models to review the design for the next version. They also found four bugs in the version that was already shipped. None of them were in the policy engine. They were all in the step that decides what a command means.
Disclosure: I drafted this article with help from an AI assistant and reviewed it myself. The bugs, the fixes and the command output below are from the project's own changelog and from running both versions.
Why the parser is the security boundary
A rule like "no rollout restarts in prod" is written against a normalized intent: action rollout-restart, resource deployment/*, namespace prod. If the parser produces a different shape for the same command, the rule doesn't match. In most guards, including this one, "no rule matched" means allow.
So a parser bug isn't a crash or a wrong answer you'd notice. It's a silent allow.
Bug 1: a flag before the subcommand
kubectl accepts global flags almost anywhere, and agents put them anywhere. This is what 0.1.5 made of a perfectly ordinary command:
$ aegis check kubectl --json -- kubectl -n prod rollout restart deploy/web
0.1.5: ALLOW rollout--n prod/*
ALLOW rollout--n restart/*
ALLOW rollout--n deployment/web (namespace lost)
0.2.1: ALLOW rollout-restart deployment/web ns=prod
The parser took -n as the rollout subcommand, then read everything after it as resources. The namespace was gone, so a namespace-scoped rule about rollouts in prod could never match. The fix takes the subcommand before merging global options, and consumes global options that appear between rollout and restart too, because kubectl accepts that placement as well.
Bug 2: impersonation was treated as noise
--as changes who the API server thinks is making the request. Rules scoped to the agent's own identity stop applying server-side, which makes the client-side hook the only place that sees the swap before it happens.
$ aegis check kubectl --json -- kubectl --as admin rollout restart deploy/web -n prod
0.1.5: ALLOW rollout---as admin/* (and two more intents like it)
0.2.1: ALLOW rollout-restart deployment/web ns=prod
BLOCK impersonate user/admin ns=prod [block-kubectl-impersonation]
In 0.2.1, --as, --as-group and --as-uid each become their own impersonate intent, and the example policy blocks them. kubectl rollout --as admin restart ... parses the same way.
Bug 3: the same object under different names
Rules are written against one canonical name per object. The parser didn't always produce it:
kubectl cordon node1 0.1.5: node1/* 0.2.1: node/node1
kubectl exec web-0 -n prod 0.1.5: web-0 0.2.1: pod/web-0
kubectl delete sc fast 0.1.5: sc/fast 0.2.1: storageclass/fast
A rule for storageclass/* missed sc/fast. A rule for node/* missed every drain and cordon.
There was a fourth case in this family that failed the other way. kubectl drain node1 --ignore-daemonsets was rejected as malformed (exit 65), because the parser thought the flag needed a value. The hook fails closed, so the agent got blocked instead of waved through. That's annoying, but it's the right way to be wrong. Failing closed turned a parser bug into friction instead of a hole.
Bug 4: letting untrusted rules stall the pipeline
This one isn't a parsing bug, but it's the one that worried me most, because it's the exact attack the project exists to stop.
Aegis only lets a rule vote if it passes two checks: nobody edited it after it was signed, and its author was allowed to write that kind of rule. A rule that fails either check is discarded. That stops a line planted in a ticket or a Slack thread from becoming policy.
But there was a path that ran before those checks. When the environment of an action couldn't be resolved, every matching rule forced ESCALATE. So an unauthorized rule scoped to env: prod could stall any action whose environment was unknown. That's a denial of service on the guardrail, from exactly the kind of rule it's meant to ignore. In 0.2.1 those rules are discarded on that path too, and reported in discarded[] like everywhere else.
What I'd tell anyone writing a command guard
These apply whether you're writing hooks for Claude Code, a wrapper script, or an admission policy.
- Normalize before you match. Write rules against one canonical form, and test that every spelling (aliases, short kinds, plural kinds, bare names) reaches it.
- Test flag placement, not just flags. Generate the same command with global options before the verb, between the verb and its subcommand, and at the end, and assert that the intent is identical.
-
Flags that change identity or target are part of the intent.
--as,--context,--kubeconfig,-n. Never file them under "options we can ignore". - Fail closed on anything you don't fully understand. An unknown option should block, not fall through to "no rule matched".
- Get someone else to read the design. None of these came from my tests. They came from reviewers reading the code while checking a design for something else.
Try it
The fixes are in 0.2.1 and later (0.3.2 is current on PyPI). If you wrote rules against the old forms (node1/*, bare pod names, short kinds), update them. docs/constraints.md has the normal forms.
pip install aegis-devops && aegis init .aegis
# Claude Code
claude plugin marketplace add moneytool/aegis-devops
claude plugin install aegis-devops@aegis-devops
# Gemini CLI
gemini extensions install https://github.com/moneytool/aegis-devops
# Codex, Copilot, VS Code, Cursor, OpenCode
aegis install codex # or copilot | vscode | cursor | opencode
The follow-up, on enforcing the same policy in AWS itself: My agent guardrail only lived on the laptop. So I compiled it into AWS.
If you've found a command shape that slips past your own agent guard, I'd like to hear it. Issues are open at github.com/moneytool/aegis-devops.
Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.