Dev.to AI 🤖 Ai 👁 0 📖 14 min read

OpenShell: Building a Security Boundary Around AI Agents

Hello, I'm Shrijith Venkatramana, and I'm building LiveReview — a blast-radius aware AI code review built for your business-critical systems. Star us to help devs discover the project, give it a try, and share your feedb

OpenShell: Building a Security Boundary Around AI Agents

Hello, I'm Shrijith Venkatramana, and I'm building LiveReview — a blast-radius aware AI code review built for your business-critical systems. Star us to help devs discover the project, give it a try, and share your feedback to help improve the product.

An AI coding agent with shell access has crossed a boundary that a chatbot never crossed.

A chatbot can suggest:

rm -rf /important/directory

An agent can actually run it.

It can read your repository, inspect environment variables, install packages, call external APIs, modify configuration, create files, launch subprocesses and keep doing all of this after you have stopped watching.

That changes the security problem.

The old question was:

“Can I trust this program?”

The more useful question for agents becomes:

“What happens when this program cannot be trusted for the next ten minutes?”

That is the problem NVIDIA OpenShell is trying to solve.

OpenShell is a runtime for autonomous AI agents. The important architectural idea is simple:

Put the security boundary outside the agent.

The agent gets a sandbox. The sandbox gets a policy. The operating system and runtime enforce that policy.

This article looks at why that distinction matters, how OpenShell works, what its policy model looks like, and where the economics and engineering trade-offs appear.

At the time of writing, the latest OpenShell release is 0.1.1, while NVIDIA's documentation still describes the software as alpha. So think of this as an architectural tour of an emerging runtime, not a production certification.

1. An AI agent is starting to look like a foreign process on your machine

Traditional software usually has a relatively stable authority model.

You install a compiler. The compiler can read your source tree.

You run a test suite. The test runner can access whatever the test process can access.

You start a server. The server has the permissions you gave its Unix user.

With an AI agent, the software is producing the next action dynamically.

A coding agent might do this:

read issue
    ↓
inspect repository
    ↓
search documentation
    ↓
install dependency
    ↓
run tests
    ↓
discover failure
    ↓
modify code
    ↓
run shell command
    ↓
call GitHub API
    ↓
spawn another process
    ↓
try another approach

The problem is not that any individual operation is unusual.

The problem is composition.

An agent can combine five individually reasonable permissions into a dangerous capability.

For example:

read ~/.config
+
read environment variables
+
execute Python
+
internet access
+
GitHub token
=
potential credential exfiltration

The model does not need to be malicious for this to matter.

A prompt injection buried inside a README can tell the agent to inspect a local file. A package installation step can execute arbitrary code. A tool response can contain instructions that compete with the original task.

This is exactly the sort of environment studied in AgentDojo, which evaluated tool-using agents against adversarial content. The benchmark contained 97 realistic tasks and 629 security test cases involving things such as email, banking and travel workflows.

The important lesson is architectural:

Prompt-level instructions are not a sufficient security boundary for a process that can act on the world.

That is why OpenShell is interesting.

It does not try to make the model perfectly trustworthy.

It assumes the opposite.

2. Sandboxing is an old operating-systems idea

OpenShell may sound like an AI-specific security invention, but its central idea is much older than neural networks.

In 1975, Jerome Saltzer and Michael Schroeder published The Protection of Information in Computer Systems. Among the principles they discussed were least privilege and fail-safe defaults.

The intuition is straightforward:

Give a program only the authority required for its job, and make access unavailable unless it has been explicitly granted.

The same philosophy appears in the browser.

When Google Chrome shipped in 2008, Adam Barth, Collin Jackson, Charlie Reis and the Chrome team described a design in which rendering happened in restricted processes while a more privileged browser process mediated access to the outside world.

The reason was practical.

If an attacker found a vulnerability in the renderer, arbitrary code execution inside the renderer should still have limited access to the machine.

That distinction is worth remembering:

Compromised process
        ≠
Compromised machine

OpenShell is applying a similar idea to agents.

You can think of it as:

AI agent
    |
    | requests
    v
sandbox boundary
    |
    +---- filesystem policy
    +---- process policy
    +---- network policy
    +---- credential policy
    |
    v
host / enterprise infrastructure

The agent is allowed to be wrong.

The security boundary is supposed to remain correct anyway.

This is a very different philosophy from:

System prompt:
"Never access secrets."

That is an instruction.

A sandbox policy is an enforcement mechanism.

3. OpenShell's mental model: the browser tab for agents

NVIDIA describes OpenShell using a useful analogy: the browser tab.

A browser tab is not allowed to behave like an arbitrary native process with unrestricted access to your computer.

It operates inside a bounded environment.

OpenShell tries to give an agent the same relationship with the host.

You can launch an agent such as Claude Code inside a sandbox:

openshell sandbox create -- claude

The architecture then separates several responsibilities.

There is a gateway, which acts as the control plane.

There is a sandbox, where the workload executes.

There is a supervisor, which lives on the trusted side of the workload boundary and mediates privileged operations.

The agent itself becomes a child process inside this environment.

Conceptually:

                     CONTROL PLANE
                          |
                       Gateway
                          |
                 policy / providers
                          |
              +-----------+-----------+
              |                       |
        Supervisor                Supervisor
              |                       |
        Sandbox A                 Sandbox B
              |                       |
        Claude Code              Codex
              |                       |
          processes              processes

The distinction between the supervisor and the agent is important.

The agent should not control the mechanism that controls the agent.

That is analogous to putting authentication checks outside the application instead of asking the application to certify that its own requests are safe.

This is one of the strongest parts of the design.

If the agent modifies its own source code, installs a new package, spawns a child process or becomes confused by prompt injection, those actions still encounter the external enforcement boundary.

OpenShell currently exposes several runtime backends, including Docker, Podman, Kubernetes and VM-backed execution. The exact isolation machinery varies, but the policy model is intended to remain stable.

4. The interesting part is the policy: an agent gets a capability graph

A useful way to think about an OpenShell policy is as a graph.

Let:

F = filesystem permissions
N = network destinations
P = executable/process identities
S = credential bindings

The agent's actual capability set is roughly a subset of:

F × N × P × S

This is not a formal security metric. It is a useful mental model.

Imagine an agent has:

3 executable identities
4 network endpoints
2 HTTP operation classes

That gives at most:

3 × 4 × 2 = 24

meaningful executable-to-endpoint-to-operation combinations to reason about.

With unrestricted outbound networking, the number of possible destinations becomes effectively unbounded from the policy's point of view.

That is the value of least privilege: reduce the graph before you have to reason about its behavior.

A simplified policy can look like this:

version: 1

filesystem_policy:
  include_workdir: true
  read_only:
    - /usr
    - /lib
    - /etc
  read_write:
    - /tmp

landlock:
  compatibility: best_effort

network_policies:
  github_api:
    name: github-api-readonly

    endpoints:
      - host: api.github.com
        port: 443
        protocol: rest
        enforcement: enforce
        access: read-only

    binaries:
      - path: /usr/bin/curl
      - path: /usr/bin/gh

This says something considerably stronger than:

"The agent should probably only use GitHub."

It says:

These binaries
    ↓
may connect to
    ↓
this endpoint
    ↓
using
    ↓
these permitted operations

OpenShell can go further and constrain HTTP methods and paths.

For example:

network_policies:

  github_repository_api:

    endpoints:
      - host: api.github.com
        port: 443
        protocol: rest
        enforcement: enforce

        rules:
          - allow:
              method: "*"
              path: "/repos/acme/project/**"

        deny_rules:
          - method: "*"
            path: "/repos/acme/project/hooks"

          - method: "*"
            path: "/repos/acme/project/hooks/**"

    binaries:
      - path: /usr/bin/gh

Now you are describing something closer to an application firewall:

binary
+
host
+
port
+
protocol
+
HTTP method
+
path

This matters because an agent does not merely need “Internet access.”

It often needs one very specific network action.

For example:

GET /repos/acme/project/issues

is a very different capability from:

POST /repos/acme/project/hooks

The second might enable an entirely different class of side effect.

This is where agent security starts looking less like prompt engineering and more like operating-system protection.

5. Under the hood, OpenShell combines several old security mechanisms

The interesting technical detail is that OpenShell does not depend on a single sandbox primitive.

It layers multiple controls.

Filesystem: Landlock

OpenShell uses Linux Landlock to constrain filesystem access.

Conceptually:

read:
    /usr
    /lib
    /etc

read/write:
    /sandbox
    /tmp

everything else:
    deny

This is implemented at the kernel level.

That matters because the agent cannot simply decide:

"Ignore the policy."

and turn it off from inside the process.

The current implementation distinguishes between static filesystem/process controls and dynamic network controls. Filesystem restrictions are established when the sandbox starts, while network policy can be changed while the sandbox is running.

Processes: reduced privilege

The workload runs as an unprivileged identity with reduced capabilities.

OpenShell's architecture explicitly aims to avoid giving the agent Linux capabilities inside the workload.

The point is defense in depth.

Suppose the agent contains a vulnerability.

Suppose the agent executes malicious code.

Suppose a dependency behaves badly.

The resulting process still has a restricted operating-system authority.

Network: deny first

This may be the most important practical control.

Outbound networking starts from a deny-by-default posture.

Connections are mediated through the OpenShell network path, where policy can inspect the destination, calling executable and, for supported protocols, higher-level HTTP operations.

That allows a policy like:

/usr/bin/gh
    -> api.github.com:443
    -> GET
    -> /repos/acme/project/**

while denying:

/usr/bin/gh
    -> random-attacker.example

and potentially:

/usr/bin/gh
    -> api.github.com
    -> POST /hooks

The current architecture also tracks executable identity using the executable's observed path and SHA-256 digest. That closes an interesting hole:

allowed binary
    ↓
replace binary
    ↓
same path

A path alone is weaker than:

path + executable identity

because otherwise an allowed executable could potentially be replaced by another program.

Credentials: give the agent access without handing it the secret

This is particularly relevant for coding agents.

Suppose your agent needs a GitHub token.

The naive approach is:

export GITHUB_TOKEN=...
claude

Now the process has the secret.

Any code it executes may potentially inspect it.

OpenShell instead has a provider model in which credentials can be stored outside the sandbox and bound to authorized endpoints.

Conceptually:

Agent
  |
  | "call GitHub"
  v
OpenShell policy
  |
  | allowed?
  v
credential binding
  |
  v
GitHub

The agent can get the capability without necessarily getting the raw credential as ordinary application data.

This is a much better abstraction for autonomous software.

The question becomes:

“Which service may this workload access?”

rather than:

“How do I smuggle the password into this process?”

6. The operational challenge is allowing autonomy without turning policy into babysitting

There is an uncomfortable tension here.

A very restrictive agent may be safe but useless.

Imagine an agent starts working and discovers it needs:

registry.npmjs.org

but the network is denied.

You have two choices.

The old approach is:

developer notices failure
    ↓
opens policy file
    ↓
adds endpoint
    ↓
restarts things
    ↓
agent continues

That becomes expensive very quickly.

OpenShell has a policy advisor workflow that turns a denied request into a proposed policy change.

The important property is that the agent can propose the permission, but the proposal does not automatically become policy.

Conceptually:

agent request
    ↓
DENY
    ↓
agent explains:
"I need registry.npmjs.org"
    ↓
policy proposal
    ↓
validation / review
    ↓
APPROVE
    ↓
live policy update
    ↓
retry

This preserves a useful separation:

agent proposes
human authorizes
runtime enforces

OpenShell also includes a policy prover. The current implementation models policy, credentials and binary capabilities and uses formal reasoning to identify changes such as newly reachable credentialed destinations.

That introduces an interesting idea:

security policy itself becomes something you can analyze mechanically.

You can think of the prover as asking questions such as:

Did this change introduce a new credentialed path?

Did this change introduce a new capability?

Did this change make a previously inaccessible endpoint reachable?

This is much more tractable than trying to prove that an LLM will never make a bad decision.

A small economics calculation

Suppose an agent generates:

20 permission requests/day

and each one takes:

20 seconds

of human attention.

Then:

20 × 20 sec = 400 sec
             ≈ 6.7 min/day

For one agent, that is manageable.

Now suppose you operate 100 agents:

100 × 20 × 20 sec
= 40,000 sec/day
≈ 11.1 hours/day

The exact numbers are hypothetical, but the operational principle is real:

the cost of agent security depends heavily on how often humans must intervene.

A policy system therefore has to optimize for two things at once:

minimal authority
+
minimal interruption

That is why live policy updates, narrow proposals, deterministic checks and audit logs are more interesting than simply putting the agent in Docker and stopping there.

Auditability matters too

OpenShell records structured security events for things such as:

allowed network access
denied network access
policy changes
security findings

That gives operations teams a way to answer:

What did the agent try to do?

What did the runtime allow?

What did the runtime deny?

Which policy allowed it?

Which executable made the request?

For autonomous systems, this starts looking like observability for authority rather than merely observability for performance.

A distributed system has traces for:

request A -> service B -> database C

An agent runtime needs something similar for:

agent -> binary -> endpoint -> credential -> action

7. What OpenShell does not solve

The sandbox boundary is powerful precisely because it has a clear scope.

It also has clear limits.

Suppose you allow:

agent -> api.github.com

The sandbox cannot infer whether the agent should create a repository release.

That is an application-level authorization problem.

Similarly, if you allow:

agent -> your production API

then the sandbox cannot magically determine whether:

DELETE /customers/123

was semantically appropriate.

There are at least four different layers of security:

1. Model behavior
2. Agent/application behavior
3. Runtime containment
4. Infrastructure authorization

OpenShell mainly strengthens layer 3 and connects it to parts of layers 2 and 4.

It does not make the model trustworthy.

It does not eliminate prompt injection.

It does not eliminate supply-chain attacks.

It does not guarantee that an allowed API call is logically correct.

It does not make a kernel vulnerability disappear.

And a sandbox is only as good as the isolation boundary beneath it.

This is another lesson from decades of systems security: containment reduces the consequences of failure; it does not make failure impossible.

The browser analogy is useful again.

Chrome's sandbox did not make browsers bug-free.

It changed the consequence of a renderer compromise.

That is the right way to think about an agent sandbox.

The agent may still go wrong.

The question is whether “the agent went wrong” means:

bad code in /tmp

or:

production database deleted
AWS credentials exfiltrated
private repositories copied

Those are very different failure domains.

The larger idea

AI agents are slowly turning software from:

programs that execute instructions

into:

programs that continuously choose instructions

Once that happens, the operating system becomes relevant again.

The shell becomes relevant.

Permissions become relevant.

Process identity becomes relevant.

Credential boundaries become relevant.

Audit logs become relevant.

Least privilege becomes relevant.

In other words, many of the ideas that operating-systems and security engineers have worked on for decades suddenly become central to AI engineering.

OpenShell is interesting because it treats the agent as a potentially untrusted process and asks a very old systems question:

What authority does this process actually need?

That is a much more scalable question than asking the model to behave itself.

For a developer, the mental shift is simple:

Before:

AI agent
   |
   +---- full machine
   +---- full filesystem
   +---- full network
   +---- credentials

After:

AI agent
   |
   v
OpenShell sandbox
   |
   +---- specific files
   +---- specific binaries
   +---- specific endpoints
   +---- specific API operations
   +---- specific credentials

The real promise of agentic software may therefore depend less on making agents perfectly aligned and more on making them containable when they are wrong.

And that is an old engineering trick.

We learned it from browsers.

We learned it from operating systems.

We are now applying it to machines that can reason about what to do next.

Conclusion

The deepest idea behind OpenShell is not Docker, Landlock, seccomp or YAML.

It is the decision to move trust outside the agent.

Once an autonomous system can execute arbitrary code, persistent memory, external tools and network calls, “please behave” stops being a sufficient security architecture.

A better model is:

Give the agent enough authority to be useful.

Put the authority behind a boundary it cannot rewrite.

Make that boundary explicit.

Make changes reviewable.

Make violations observable.

That is the basic architecture required when software starts acting more like an autonomous operator than a traditional program.

The interesting question for developers is no longer whether agents will get shell access.

They already are.

The interesting question is: what is the smallest world you are willing to let an agent operate in?


Your team's attention is limited, and the deluge of AI-generated code is making it harder to keep production reliable and secure without slowing you down.

I'm building LiveReview, a blast-radius aware AI code review built for your business-critical systems.

Instead of presenting every diff with equal emphasis, LiveReview scores each change by blast radius — how far its impact reaches through your call graph — so you can focus attention where it actually matters.

Spend code review effort where business risk is highest — not spread evenly across every diff.

⭐ Star it on GitHub:

GitHub logo HexmosTech / LiveReview

Blast-Radius Aware AI Code Review for Business-Critical Systems

LiveReview

gitleaks.yml osv-scanner.yml govulncheck.yml semgrep.yml dependabot-enabled mcp-testcases.yml

LiveReview: Blast-Radius Aware AI Code Review for Business-Critical Systems

LiveReview is an AI code reviewer that scores every hunk of a diff by blast radius: how far a change reaches through your call graph, how much persistent state it touches, and how well-tested it is. A 3-line change to a shared auth check can outrank a 300-line UI tweak. Your team's attention goes to the highest-risk code first, not spread evenly across every diff.

blast-radius-demo.mp4

LiveReview's Blast Radius & Review Priority scoring, live in the diff viewer.
















The exact math, not a black box Visualize blast radius at a glance Every factor that feeds the score

How does Blast Radius scoring work? (a more technical explanation)

Here's the goal:

  • A 3-line fix in a function used by 40 other files, that also writes to a database, should score high.
  • A 300-line UI change in one file, fully covered by…




Click below to try LiveReview with your codebase:

LiveReview Banner

📰 Read the original article on Dev.to AI

Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.