Dev.to AI πŸ€– Ai πŸ‘ 0 πŸ“– 4 min read

Claude Code Mods Aren't Sandboxed. Neither Is Your Understanding of "Sandboxed".

Two statements about Claude Code mods, both true, both from Anthropic's own docs: The module runs in a sandbox of its own, with no DOM and no Node. "Mods aren't sandboxed." The first two days of the mods launch have

Two statements about Claude Code mods, both true, both from Anthropic's own docs:

  1. The module runs in a sandbox of its own, with no DOM and no Node.
  2. "Mods aren't sandboxed."

The first two days of the mods launch have been one long argument about which of these is the lie. Neither is. The confusion comes from the word "sandbox" doing two different jobs, and once you separate them, the real trust boundary snaps into focus. It is not where most people think it is.

Sandbox #1: the JavaScript runtime

When you write a mod, your module gets a private little world. No document. No window. No Node globals. If your code wants to touch anything outside itself, files, processes, the network, the UI, it goes through one object: $.

export function register(on) {
  on("tool.call", { tool: "Bash" }, async ($, e, next) => {
    const listing = await $.fs.readDir("/tmp");   // through $, not fs
    return next(e);
  });
}

This is sandbox #1. It is a language-level sandbox: it constrains how your code reaches the outside world, not whether it can. Every capability lives behind the $ API, which means the engine can see every request, log it, and (in theory) gate it. Think of it like a phone where every app must use the official APIs. The APIs still include the camera.

Sandbox #2: the one that does not exist

Here is what the documentation says, and I am quoting because paraphrasing would soften it:

"A mod is code that runs with your permissions. It can read and write your files, start processes, and make network requests. Install mods only from authors and marketplaces you trust."

And:

"Mods aren't sandboxed. If you turn on sandboxing, the sandbox isolates the Bash commands Claude runs, and a process that a mod starts runs outside it."

Read that second one twice. Claude Code has a sandboxing feature, and it sandboxes Claude's Bash commands. A mod that starts its own process steps cleanly outside of it. The fence was built around the agent, not around the extension.

The capability list gets worse the further you read. A mod can:

  • Read your secrets: environment variables and settings files, "including an API key you keep in either".
  • Approve tool calls you blocked: a mod that approves tool calls can green-light one "that an ask rule would prompt for, or that one of your own PreToolUse hooks blocked". Your hook said no. The mod says yes. The mod wins.
  • Rewrite events before you see them: hooks form a chain, and a mod can observe, rewrite, or fully answer any event. Including the ones your audit logger was counting on.

So when someone says "mods are sandboxed", they are describing the JS runtime. When Anthropic says "mods aren't sandboxed", they are describing your files, your processes, your network, and your API keys. Both true. The second one is the one that matters.

The one hard boundary

There is exactly one thing Anthropic drew a hard line around, and it is telling:

"A mod can restyle much of Claude Code's interface, but not the permission prompt. It can't change what a prompt shows you."

A mod can redraw almost the entire UI, panes, bands above the prompt, toasts, but the permission dialog stays Claude Code's own. Whatever else a mod draws, the moment that asks "are you sure?" cannot be faked, restyled, or have its contents swapped.

Think about why that specific line exists. If a mod could restyle the permission prompt, it could show you "run tests?" while actually approving rm -rf. Anthropic hardened the one UI element where your eyes are the security control, and left everything else open. That tells you exactly what threat model they designed for: the mod is untrusted code with your privileges, and the permission prompt is the last honest surface in the room.

The mental model that actually works

Stop picturing a browser iframe. The right mental model is much older and much simpler:

Installing a mod is giving someone your shell.

Not a restricted shell. Not a shell with an auditor watching. Your shell, with your env vars, your files, your network, and the ability to overrule the guardrails you built for the agent. The JS-runtime sandbox is real engineering, it keeps modules from stepping on each other and gives the engine a clean interception point, but it was never a security boundary between the mod and your machine.

Once you have that model, the hygiene checklist writes itself:

  1. claude plugin validate before you install, not after. It lists the events a mod handles and what it asks Claude Code to do, file reads, network requests, without running any code. Read it like you would read the permissions screen on a phone app. If a context-bar mod wants network access, ask why.
  2. Know your off switches. Disable one plugin from the Installed tab in /plugin. Start a session with --safe-mode to drop all customizations. Set "disableAllHooks": true in ~/.claude/settings.json for the wide kill switch. Organizations get allowManagedModsOnly, which stops user-installed mods from loading while skills, commands, and MCP servers keep working.
  3. Audit the small ones hardest. A 40-line UI mod feels safe because it is small. But 40 lines is enough to read ~/.claude/settings.json, exfiltrate an API key, and approve the tool call that covers its tracks. Size is not a security property.

None of this means you should not install mods. It means you should install them the way you would hand your laptop to a colleague: only to people you trust, and knowing exactly what they can reach while they have it.

So here is my question: what does your claude plugin validate say about the mods you have already installed? And did you read it before, or just now?

πŸ“° Read the original article on Dev.to AI

Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes β€” full credit and traffic to the original publisher.