Dev.to AI 🤖 Ai 👁 0 📖 5 min read

AWS Lambda MicroVMs Run Your Agents. Here Are 17 Invariants to Verify Before You Deploy. With 8 Hands-On Labs.

✓ Human-authored analysis; AI used for formatting and proofreading. AWS Lambda MicroVMs put your code including coding agents and AI assistants inside a Firecracker microVM that boots in ~125–150 ms. "Coding agent host

✓ Human-authored analysis; AI used for formatting and proofreading.

AWS Lambda MicroVMs put your code including coding agents and AI assistants inside a Firecracker microVM that boots in ~125–150 ms. "Coding agent hosts" is an explicit use case. So the question for anyone deploying an agent into one: what must be true about this microVM before I trust it with my account?

Stave now answers that. The built-in catalog gained 16 Lambda MicroVM controls (CTL.LAMBDA.MICROVM.*), plus a graph-reachability check for the one risk per-resource scanners structurally cannot see. And there are 8 hands-on labs one per scenario, each self-contained that you can follow in an AWS sandbox from broken to fixed.

This post is about: what's new, the control families at a glance, the lab methodology, and a short review of each lab.

What's new

  1. 16 built-in MicroVM controls. Rebuild Stave and they ship in the catalog (now 2,690 controls). There are no custom rules to write. They cover the network, identity, IAM, supply chain, runtime, and snapshot surfaces of a MicroVM.
  2. The compound-path check (graph reachability). A microVM's execution role can reach a production secret through an assume-role hop where every individual hop passes its own check. That's not a per-resource control. It's reachability over a graph. Stave models the edges from the captured IAM policies and computes it with two independent engines, Soufflé (Datalog) and Z3 (SMT), that agree.
  3. A self-contained lab methodology. Each lab plants one misconfiguration, captures it, projects it into an obs.v0.1 snapshot, evaluates it, then remediates and re-verifies to zero findings. Every command and output is verified against the real binary.

The control families

Each control reads a derived signal from your captured AWS configuration:

Surface Controls What they verify
Network approved subnets, restricted SG ingress, connector role the microVM's network connector reaches only approved private subnets with restricted ingress, via a least-privilege ENI role
Identity ingress auth, trust-policy TagSession, connector origin, workload identity claims ingress is authenticated, the execution role's trust policy supports session tagging, and the workload has governable identity claims
IAM roles execution role, build role three separate roles (execution, build, network) are least-privilege — no wildcards, no iam:PassRole
Supply chain approved base image, S3 artifact bucket public access, versioning the image is built from an approved base, and the S3 artifact (your Dockerfile + code) is private and versioned
Runtime entropy reinit, idle limit, runtime limit entropy is reinitialized on resume, and idle/runtime durations stay within your org limit (not the 8-hour AWS ceiling)
Snapshot secrets-in-snapshot no secret is loaded at init and baked into the Firecracker memory snapshot
Compound path reachability no execution role transitively reaches a secret through an assume-role chain

Two of these are:

  • Secrets in the snapshot. MicroVMs boot fast by resuming from a memory snapshot taken after init. Load a secret at module scope and it's in that snapshot in every microVM launched from the image, on every resume, indefinitely. No per-resource cloud scanner looks inside a Firecracker snapshot.
  • The compound path. The execution role only has sts:AssumeRole on one role. The intermediate role only has secretsmanager:GetSecretValue on one secret. Both pass every per-resource check. The composition role → assume → role → read → secret is lethal, and only graph reachability sees it.

How the labs work

Stave evaluates a snapshot, not a live account. The loop in every lab:

capture (aws → jq) → obs.v0.1 → stave apply → finding
                                      │
                              remediate the resource
                                      │
        re-capture → obs.v0.1 → stave apply → 0 findings

A small transform defaults every un-captured resource to compliant, so each lab captures only the one resource it's testing. That makes them self-contained where no lab depends on another's setup. Per-resource controls run through Stave's CEL engine; the compound path runs through Soufflé and Z3.

The 8 labs

Lab 1 — Baseline: A Compliant MicroVM Orientation. Build a compliant execution role, capture it, and watch all 16 controls pass with exit code 0. You learn the obs.v0.1 shape and the capture → transform → evaluate loop you'll reuse everywhere.

Lab 2 — Over-privileged Execution Role Plant a role with s3:*, secretsmanager:*, and iam:PassRole, and a trust policy missing sts:TagSession. Two controls fire (over-privilege and TagSession are independent invariants). Scope to least privilege, re-verify to zero.

Lab 3 — Trust Policy Missing sts:TagSession The same TagSession defect in isolation: a least-privilege role whose trust policy just lacks session tagging. Exactly one control fires. This demonstrates that Stave reports the one thing that's wrong, without noise.

Lab 3B — Trust Policy Over-Broad sts:TagSession The mirror of Lab 3: sts:TagSession is present, but granted through an sts:* wildcard. A separate control (TAGSESSION.WILDCARD) fires on the over-broad grant; scope the actions and re-verify. "Has the permission" and "grants it correctly" are two distinct invariants, so they're two controls that never fire together.

Lab 4 — Compound Path: Graph Reachability The differentiator. The per-resource controls pass every role, yet the execution role reaches a production secret through an assume-role hop. Soufflé and Z3 both detect it (sat with a witness); break one edge and both report no path (unsat). This is the lab to show a skeptic.

Lab 5 — Public Artifact Bucket The S3 bucket holding your image's source has Block Public Access off and versioning disabled with source-code exposure plus no deployment audit trail. Two supply-chain controls fire; enable PAB + versioning, re-verify to zero.

Lab 6 — Excessive Idle and Runtime An agent that idles or runs for 8 hours unsupervised is a cost sink and a risk window. The idle/runtime limits aren't in any AWS API response, so they're asserted from how the microVM was launched. A good lesson in signals that don't come from a describe call. Lower within the org limit, re-verify.

Lab 7 — Secrets Baked Into the Snapshot An image whose init loads a secret at module scope, baked into the memory snapshot. One control fires. The fix is structural: defer secret loading to a runtime lifecycle hook so it runs after the snapshot.

Why this matters for agents

An agent inside a microVM is a non-human identity with the execution role's full power, a network path, a supply chain, and a memory snapshot. Every one of those is an attack surface, and most of them are invisible to "scan the deployed config for a public flag." The compound path especially: as agents chain tool calls across roles, the lethal capability is rarely on any single resource. It's in the reachability.

Declaring these as invariants and verifying them on every configuration change, deterministically, with explainable findings is how you keep "the agent host is secure" from being one team's opinion against another's.

Get started

  1. Build Stave (cd stave && make build). The MicroVM controls ship in the catalog.
  2. Open an AWS sandbox (not production) with Lambda MicroVM access.
  3. Start with Lab 1, then work through 2–7. Each ends at zero findings, so you always know you're back to a clean state.

The labs are short, self-contained, and every command is verified. If you're deploying agents into MicroVMs, run them before you deploy.

📰 Read the original article on Dev.to AI

Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.