Six Controls from Nine AI Agent Incidents
An AI agent runs with your credentials and acts before you can stop it. GitGuardian analyzed nine public incidents between May 2025 and July 2026 to reverse-engineer what breaks in production. The pattern is consistent:
An AI agent runs with your credentials and acts before you can stop it. GitGuardian analyzed nine public incidents between May 2025 and July 2026 to reverse-engineer what breaks in production. The pattern is consistent: agents leak credentials, exceed spending caps through parallel tool calls, and execute instructions embedded in untrusted input.
The post-mortem analysis maps real failures to six concrete controls. This is not a threat model. This is what broke when agents met production workloads.
Three Attack Surfaces
An agent exposes three distinct surfaces:
Host surface: Everything the process can reach. Filesystem, process execution, network, environment variables, hardware devices.
Identity surface: The account the agent acts as, with its privileges and scopes. Every credential it can discover along the way.
Agent surface: The part with no equivalent in a traditional process. The agent's configuration, its tool registry, its reasoning loop, and the workspace it reads to make decisions.
Of the nine incidents, the agent surface appeared in five, always combined with the host or identity surface. An attacker who can modify the agent's configuration or inject instructions into its workspace can pivot to credentials or execution.
Incident Patterns
The nine incidents break into three categories. The six controls that follow are derived specifically from these failure modes, not from theoretical threat modeling.
Credential leaks: Agent reads untrusted input (GitHub issue, email, settings file, web page) containing instructions. The instruction opens the door. The credential is what the attacker came for.
Runaway costs: Agent spawns parallel tool calls that each stay under individual limits but collectively exceed budget. No single call triggers a spending cap, but the aggregate cost blows through monthly allocation.
Tool misuse: Agent invokes a tool with parameters the developer did not anticipate. The tool has legitimate access, but the agent's reasoning produces a call sequence that deletes data, modifies configuration, or exfiltrates files.
Six Controls
GitGuardian derived six controls from the incident analysis. None of them prevent prompt injection. They limit blast radius.
1. Sandbox the Agent
Run the agent in a container or VM with no access to the host filesystem, no process execution, and network egress limited to approved endpoints. The sandbox boundary is the first failure point.
Implementation shape: Docker container with read-only root filesystem, no privileged mode, and a network policy that denies all egress except to specific API endpoints. Use a sidecar proxy to enforce the allow list.
Failure mode: Agent needs to read a file from the host to complete a task. Developer mounts the host filesystem into the container. The sandbox is now a suggestion.
2. Scope Credentials
Give the agent the minimum credential scope required for its task. Do not give it a credential that can create other credentials. Do not give it a credential that can modify its own permissions.
Implementation shape: OAuth token with a scope limited to read-only access to a single repository. Service account with IAM permissions that allow only the specific API calls the agent needs. Rotate the credential every 24 hours.
Failure mode: Agent needs to perform a task that requires a broader scope. Developer upgrades the credential to admin. The agent now has the keys to everything.
3. Lock Configuration
Store the agent's configuration outside its workspace. Do not let the agent read or modify its own tool registry, system prompt, or execution policy.
Implementation shape: Configuration lives in a separate repository or secret store. The agent process has read-only access to a mounted config file. Changes require a deployment.
Failure mode: Agent reads a file that contains a tool definition. The file is in the workspace because a user uploaded it. The agent registers the tool and calls it.
4. Log Every Call
Record every tool invocation with parameters, response, and the reasoning that led to the call. Store logs in a write-only destination the agent cannot access.
Implementation shape: Structured logs to a centralized logging service. Each log entry includes the tool name, input parameters, output, execution time, and the portion of the agent's reasoning trace that triggered the call. Use a log shipper that buffers locally and flushes every 10 seconds.
Failure mode: Agent makes 1,000 tool calls in parallel. The log shipper buffers locally. The buffer fills. The shipper drops logs. You have no record of what the agent did.
5. Scan the Workspace First
Run a secret scanner over the agent's workspace before the agent reads any file. Block execution if the scanner finds a credential.
Implementation shape: Pre-execution hook that runs GitGuardian or TruffleHog over the workspace. If the scanner exits non-zero, the agent does not start. Latency cost is 2-5 seconds for a workspace with 1,000 files.
Failure mode: Attacker embeds a credential in a binary file. The scanner does not parse the binary format. The agent reads the file and leaks the credential in a tool call.
6. Deny by Default
Start with a policy that denies all tool calls. Explicitly allow only the tools the agent needs. Require human approval for any tool not on the allow list.
Implementation shape: Policy engine that intercepts every tool call. The engine checks the tool name against an allow list. If the tool is not on the list, the engine sends a notification to a human operator and blocks the call until the operator approves or denies.
Failure mode: Agent needs to call a tool that is not on the allow list. The operator is asleep. The agent blocks for 8 hours. The task does not complete.
Control Trade-offs
| Control | Latency Cost | Operational Overhead | Bypass Risk |
|---|---|---|---|
| Sandbox | Low (container startup: 1-2s) | Medium (network policy maintenance) | High (filesystem mounts, privileged) |
| Scoped credentials | None | High (rotation, scope management) | Medium (scope creep, reuse) |
| Locked configuration | None | Low (config deployment pipeline) | Medium (workspace file parsing) |
| Log every call | Medium (2-5s buffer flush) | Medium (log storage, retention) | High (buffer overflow, dropped logs) |
| Scan workspace first | Medium (2-5s scan time) | Low (scanner updates) | High (binary files, obfuscation) |
| Deny by default | High (human approval latency) | High (operator availability, policy) | Low (explicit allow list) |
Spending Cap Enforcement
The runaway cost incidents expose a gap in spending enforcement. An agent can spawn parallel tool calls that each stay under individual limits but collectively exceed budget. This is the core plumbing question: how do you enforce spending caps when an agent can spawn parallel tool calls that each stay under individual limits but collectively exceed budget?
The problem: You set a $100 per-call limit. The agent makes 50 parallel calls, each costing $80. Total cost: $4,000. No single call triggered the limit.
The fix: Track aggregate spending across all active calls. Use a shared counter that increments before each call and decrements after the response. If the counter plus the estimated call cost exceeds the budget, block the call.
Implementation shape:
import threading
class BudgetExceededError(Exception):
pass
class SpendingTracker:
"""
Tracks aggregate spending across parallel tool calls.
Note: This implementation assumes single-process execution.
For distributed agents across multiple processes or machines,
use a distributed lock (Redis, etcd) or a centralized spending
service with atomic increment/decrement operations.
"""
def __init__(self, budget_cents):
self.budget_cents = budget_cents
self.active_spend_cents = 0
# Lock ensures atomic read-modify-write for active_spend_cents
self.lock = threading.Lock()
def reserve(self, estimated_cost_cents):
with self.lock:
if self.active_spend_cents + estimated_cost_cents > self.budget_cents:
raise BudgetExceededError(
f"Would exceed budget: {self.active_spend_cents + estimated_cost_cents} > {self.budget_cents}"
)
self.active_spend_cents += estimated_cost_cents
def release(self, actual_cost_cents, estimated_cost_cents):
with self.lock:
self.active_spend_cents -= estimated_cost_cents
# Track actual vs estimated for future estimates
Failure mode: The agent makes a call that costs more than estimated. The counter releases the estimated cost, not the actual cost. The tracker drifts. After 100 calls, the tracker thinks you have spent $5,000 but you have spent $8,000.
The fix for the fix: Release the actual cost, not the estimated cost. Track the delta between estimated and actual. Use the delta to improve future estimates.
Drift and Recovery
Estimation error accumulates. After 100 calls with 10% estimation error, the tracker reports $5,000 spent but actual spend is $5,500. The next call reserves $80 but costs $88. The tracker now underreports by $88. After 1,000 calls, underreporting reaches $880.
The state-management boundary is between reservation and release. The tracker must maintain two counters: reserved spend (what the agent thinks it will spend) and actual spend (what the agent has spent). The delta between them is the drift.
Detection: Compare the tracker's active spend counter to the actual spend reported by the tool provider's API. If the delta exceeds 5% of the budget, trigger an alert.
Recovery: Reset the tracker's active spend counter to match the actual spend. Cancel all pending reservations. The agent must re-reserve before making new calls. This introduces a latency spike (50-200ms per call while reservations rebuild) but prevents runaway drift.
Concrete scenario: You run an agent with a $10,000 monthly budget. After 500 calls, the tracker reports $4,800 reserved but the provider API reports $5,280 actual spend. The delta is $480 (4.8% of budget). You reset the tracker to $5,280. The next 10 calls each reserve $80. The tracker now shows $6,080 reserved. The provider API shows $6,160 actual. The delta is $80 (0.8% of budget). The drift is under control.
Failure mode for the recovery: The provider API has a 5-minute lag. You reset the tracker based on stale data. The tracker underreports by the amount of spend that occurred in the last 5 minutes. If the agent makes 100 calls in those 5 minutes, each costing $80, the tracker underreports by $8,000.
The fix for the recovery: Use the provider's real-time spend API if available. If the API has lag, add a safety margin to the reset value. If the provider reports $5,280 and the API has 5-minute lag, reset the tracker to $5,280 plus 10% ($5,808). The agent will hit the budget cap earlier than necessary, but it will not exceed the actual budget.
Credential State Management
The incidents show that credential leaks happen when secrets live in the wrong place. The question is where the boundary sits between the agent's credential store and the tools it invokes. This is the second core plumbing question: what's the state-management boundary between an agent's credential store and the tools it invokes?
Three options:
Secrets in context: The agent's reasoning loop has access to credentials. It passes them to tools as parameters. The credential appears in logs, traces, and error messages.
Secrets in environment variables: The agent process has access to environment variables. Tools read credentials from the environment. The credential does not appear in logs, but it is visible to any code the agent executes.
Secrets in a vault with per-call authorization: The agent requests a credential from a vault for each tool call. The vault checks the tool name, the agent's identity, and the call parameters before returning the credential. The credential never enters the agent's memory.
The incidents favor option 3. Every credential leak involved a credential that lived in the agent's context or environment. None involved a credential that required per-call authorization.
Implementation shape: The agent calls a vault API with the tool name and parameters. The vault checks a policy that maps (agent_id, tool_name, parameters) to a credential. If the policy allows the call, the vault returns a short-lived token. The agent passes the token to the tool. The token expires after the call completes.
The boundary is the vault API. The agent never holds the credential. The vault enforces the policy. Detection latency is the vault response time plus revocation latency.
Latency cost: 50-100ms per call for the vault round trip. For an agent that makes 10 tool calls per task, that is 500ms to 1 second of added latency.
Scaling with parallel calls: If the agent makes 50 parallel tool calls, each call adds 50-100ms. The calls run in parallel, so the total latency is still 50-100ms, not 2.5-5 seconds. The vault must handle 50 concurrent requests. If the vault has a rate limit of 100 requests per second, 50 parallel calls consume half the quota.
Authorization Failure Detection
When a vault denies a call, you need to detect the denial without leaking the policy. The vault should return a generic error (e.g., "authorization denied") without specifying which part of the policy failed. Logging the denial requires care: you cannot log the credential the agent requested, but you can log the tool name, agent ID, and a hash of the parameters.
Detection latency: The vault response time (50-100ms) plus the time to log the denial (10-20ms). Total: 60-120ms. This is the latency added to the execution loop when a denial occurs.
Concrete scenario: The agent requests a credential for a tool called delete_repository. The vault policy allows the agent to call read_repository but not delete_repository. The vault denies the request. The vault logs: {"agent_id": "agent-123", "tool": "delete_repository", "params_hash": "a3f8b9c2", "denied": true}. The agent receives a generic error. The operator sees the log and updates the policy or investigates why the agent attempted the call.
Failure mode: The vault denies 100 calls in 10 seconds. The log shipper buffers the denials locally. The buffer fills. The shipper drops logs. You have no record of the denials.
The fix: Use a separate log stream for authorization denials. Set a higher priority for denial logs. If the buffer fills, drop non-denial logs first. Monitor the denial rate. If the rate exceeds 10 per second, trigger an alert and pause new agent tasks.
Detection Without Inline Scanning
When an agent leaks a credential through a tool output, you need to detect it without parsing every response. Inline secret scanning adds latency to the execution loop. This is the third core plumbing question: when an agent leaks a credential through a tool output, how do you detect it without parsing every response, and what's the latency cost of inline scanning in the execution loop?
The problem: You run a secret scanner over every tool response. The scanner takes 200ms. The agent makes 10 tool calls per task. You have added 2 seconds of latency.
The fix: Scan asynchronously. Write tool responses to a queue. A separate process reads from the queue and scans each response. If the scanner finds a credential, it revokes the credential and alerts the operator.
Implementation shape: Tool calls write responses to a Kafka topic. A consumer reads from the topic and runs GitGuardian over each message. The consumer maintains a cache of known credentials. If the scanner finds a credential in the cache, it calls the credential provider's revocation API.
Failure mode: The queue fills faster than the consumer can process. The consumer falls behind. A credential leaks at 10:00 AM. The scanner detects it at 10:15 AM. The attacker has 15 minutes to use the credential.
The fix for the fix: Run multiple consumers in parallel. Use a partitioned topic so each consumer processes a subset of responses. Monitor consumer lag. If lag exceeds 1 minute, scale up the consumer pool.
Consumer Lag Monitoring
Consumer lag is the difference between the current offset (the latest message in the topic) and the committed offset (the latest message the consumer has processed). If the current offset is 10,000 and the committed offset is 9,500, the lag is 500 messages.
Measurement: Kafka exposes lag metrics via JMX or the consumer group API. Poll the API every 10 seconds. Calculate lag as current_offset - committed_offset.
Scaling threshold: If lag exceeds 1 minute of messages (e.g., 600 messages at 10 messages per second), scale the consumer pool by 50%. If lag exceeds 5 minutes (3,000
Originally published by Dev.to Security. Aggregated on AIWithGhost for educational purposes β full credit and traffic to the original publisher.