Sidekick part 4: safety as honestly labeled ceilings
Sidekick Part 4: safety as a stack of ceilings, each honestly labeled Most agent safety I've seen is a paragraph in the system prompt: "be careful with destructive commands." Small models ignore system prompts — we est
Sidekick Part 4: safety as a stack of ceilings, each honestly labeled
Most agent safety I've seen is a paragraph in the system prompt: "be careful with destructive commands." Small models ignore system prompts — we established that in Part 2. So Sidekick's safety model assumes the model will disobey and enforces the boundaries in code instead. This is Part 4: approvals, hard refusals, egress control, and the audit ledger — plus the ceilings, stated rather than hidden.
Layer 1: approval sets, enforced in the dispatcher
Reads auto-run. Everything else sits in APPROVAL_TOOLS — writes, deletes, general shell — gated behind inline [y/N] prompts, with /yolo, /confirm, and /readonly modes. Plan mode and read-only mode go further with deny sets (PLAN_DENIED_TOOLS, READONLY_DENIED_TOOLS) enforced at dispatch. The comment in the code says the quiet part out loud: hard gates exist because "prompt text alone does not stop disobedient models." Even exec, the read-only shell, dies in these modes — inventory commands have no place in a mode that promises zero side effects.
Layer 2: the shell denylist that assumes deception
SHELL_BLOCK_PATTERNS in src/sk/tools/shell.py reads like a museum of bad days: recursive rm at /, ~, $HOME, bare system roots — matched after quote-stripping, so rm "-rf" / doesn't slip through. mkfs, dd to devices, fork bombs by structural shape, chmod -R 777 /, curl … | sh. And crucially, patterns match against de-obfuscated variants, never the raw string alone — because attackers obfuscate and careless users quote.
My favorite entry: find / -delete. It destroys trees without any rm flag, so it sailed past the rm pattern for a while. The comment admits it: already blocked in exec, but the shell tool had no coverage. Every gap found becomes a pattern; every pattern carries the story of the gap.
Layer 3: egress deny-by-default
src/sk/egress.py splits network traffic into operator-configured (your model endpoint, MCP servers you set up — always allowed) and model-chosen (a page the model decided to read, a URL a provider handed back — the destinations an attacker reaches through prompt injection). The second category is deny-by-default: empty allowlist means no network, wildcards like * are refused at parse time, and every decision lands in the ledger.
And then the ceiling, in the module docstring: this covers fetches through the agent's tools. A shell command running curl is not covered — "a shell string denylist cannot be a network boundary." The README states it explicitly so the control is not oversold. I respect this enormously: most projects hide their ceilings; this one documents them with issue numbers (#288, #329).
Layer 4: trust fences + the ledger
src/sk/trust.py funnels five independent injection paths (repo docs, web text, search results, @path inlines, MCP results) through one chokepoint — because a fence applied at only some paths "is not a boundary." Markers are deliberately stable rather than random nonces, since rotating markers would destroy the prompt-cache discipline the test suite locks in; invisible-character smuggling is handled by a sanitizer instead. Same honesty here too: "prompt framing is advisory, not a control" — fences raise the floor; deny rules give the guarantee.
Everything lands in sk audit: tool runs, approve/deny decisions, local-vs-egress rows. The ledger is the point where safety becomes proof — the same trail from Part 1's privacy claim.
What broke
The find -delete gap above. The shell-curl ceiling, which no pattern layer can fix and which is therefore documented instead of patched. And the general lesson across all four parts of this series: each layer exists because the layer above it failed in production — instructions failed, so grounding; refusals failed, so injection; per-tool prompts nagged, so plan review up front. Safety here is sedimentary. That's the whole series: overview, grounding, voice, safety. Memory, skills, and the daemon are future parts if you want them.
Built by the Sidekick community. Repo: https://github.com/Faisal-Fayaz/sidekick
Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.