How I built a 4-tier approval gate for my AI companion: a heuristic that can never say no
The moment I connected Gmail to my desktop AI companion, I realised my approval system had a hole in it. The old rule was binary: a tool marked readOnlyHint: true skipped the gate, everything else asked β and a trusted s
The moment I connected Gmail to my desktop AI companion, I realised my approval system had a hole in it. The old rule was binary: a tool marked readOnlyHint: true skipped the gate, everything else asked β and a trusted server skipped the gate entirely. Composio was trusted, which meant a connected Gmail could send mail with zero prompts. Trusted meant "the connection was vetted", but the gate treated it as "its actions may run unreviewed". That's the hole, and fixing it meant redesigning approvals from scratch.
Ankita is my open-source (MIT), local-first desktop companion β a little companion, a lot more possible. It runs tools on your actual machine: browsers, MCP servers, shell commands. So the approval gate isn't a nice-to-have, it's the whole trust model. Here's what I built.
Four tiers, not two
src/integrations/mcp-tiers.mjs replaces the binary rule with four explicit tiers:
- Tier 0 β auto-allow: read-only tools. Looking costs nothing.
- Tier 1 β ask once: the default for an unclassified tool.
- Tier 2 β always ask: destructive verbs β send, delete, publish, pay.
- Tier 3 β deny: reachable only through an explicit blocklist entry.
Resolution order is strict and inspectable: blocklist β per-tool rule β per-app rule β app defaults β heuristic. The critical property: Tier 3 is only reachable via the blocklist. That means the heuristic can never silently deny. It can escalate a prompt, never issue a refusal.
This is a deliberate asymmetry. The destructive-verb matcher is deliberately crude:
export const DESTRUCTIVE_VERBS = /(send|delete|publish|pay|transfer|remove|destroy|revoke)/i;
Substring matching, no word boundaries β so payment and resend match too. The comment in the code says it plainly: a false positive only costs an extra prompt, and asking is already the default, so there is no reason to risk a false negative with word boundaries. A wrongly-asked question is a minor annoyance; a wrongly-sent email is irreversible. The error budget goes entirely toward asking too much.
The gate polices itself
The part I'm proudest of is lowersProtection():
/** Tier changes that lower protection are themselves approval-worthy. */
lowersProtection(serverId, previousTier, nextTier) { ... }
The companion can suggest tier changes β "Gmail keeps asking me, can I set it to auto-allow?" β but any change that lowers protection is itself an approval-worthy event. The agent cannot quietly loosen its own leash. Loosening up needs your eyes; tightening down never does.
There's a related trap in the old code worth naming: the trusted flag on an MCP server. tierFor() in mcp-manager.mjs carries this comment: "trusted is deliberately NOT consulted: it means the server connection was vetted by Ankita, not that its actions may run unreviewed." Trust in the connection is not trust in every action it might take. Every call still resolves its own tier.
Meta-tools can't hide the real action
Composio exposes meta-tools like COMPOSIO_MULTI_EXECUTE_TOOL whose arguments name the real action. Classifying by tool name would call that "one tool" and miss that it contains a Gmail send. So enclosedActions() parses the arguments, splits each slug (GMAIL_SEND_EMAIL β { app: "gmail", action: "send email" }), and the tier that matters is the worst enclosed one. And what the user sees is never the meta-tool name β describeCall() renders "gmail: send email", because the plan is explicit that COMPOSIO_MULTI_EXECUTE_TOOL tells the user nothing while "gmail: send email" tells them what they're approving.
Mail gets special treatment as the highest-regret action in the default toolset: new Composio connections start with Gmail pinned at "always ask" instead of inheriting the generic default. Reads stay harmless.
"Why did it send that mail?" must be answerable
Every decision goes through recordDecision(), which keeps per-server counters of allowed / asked / denied. When something surprising happens, the first question is "what did the gate decide, and why" β and the audit trail answers it instead of a shrug.
And on the desktop side, pending approvals live in an ApprovalRegistry (desktop/electron/approvals.mjs) that stores each request on entry β not just the emitted event β so a late subscriber like the companion island can list what's still pending and render it. The gate and the UI read the same source of truth.
All of this is pinned by tests: the tier resolution suite is 14/14 green, covering the worst-enclosed-tier rule, the blocklist-outranks-everything path, and the read-only discovery shortcut.
The design decision I keep coming back to: we made the heuristic one-directional on purpose β it can escalate toward a prompt but never issue a denial. Reviewability over coverage. A heuristic with denial power would catch more risky calls silently, at the price of refusals nobody reviewed and nobody can audit their way out of.
I'm genuinely unsure where the right line is. If you were building the approval gate for your own agent, would you ever let the heuristic deny β and what would you require before trusting it?
The code is all in the repo if you want to poke at it: https://github.com/akyourowngames/A.N.K.I.T.A β src/integrations/mcp-tiers.mjs is 286 lines, no dependencies beyond a shared file helper. I'd love to hear how you'd draw the line differently.
Originally published by Dev.to Security. Aggregated on AIWithGhost for educational purposes β full credit and traffic to the original publisher.