Dev.to AI 🤖 Ai 👁 0 📖 4 min read

What should an AI agent do when it hits its budget?

Give an agent a budget and sooner or later it reaches the edge of it. A research agent has spent 4.20 of its 5.00 USD cap, and the next step (one more paid API call, a dataset, a GPU minute) costs 1.50. What should happe

What should an AI agent do when it hits its budget?

Give an agent a budget and sooner or later it reaches the edge of it. A research agent has spent 4.20 of its 5.00 USD cap, and the next step (one more paid API call, a dataset, a GPU minute) costs 1.50. What should happen now?

This post looks at three options and shows one way to build the third, with a small open-source example you can run offline.

Disclosure: I'm with send21. The example uses send21 to prepare a payment request, but the pattern works with any approval channel.

Over-budget handoff in mock mode

The problem

Spend caps exist because agents are not good judges of when something is worth paying for. A loop, a bad plan, or a prompt injection can turn "one more call" into a large bill. So we put a hard number in front of the agent.

The cap is the easy part. The hard part is what the agent does when the cap says no, because that is exactly the moment when the task is often nearly done and the extra spend is often justified.

Option 1: raise an error

The simplest thing: throw, log, stop.

It is safe. Nothing gets spent. But the run is lost, the context is gone, and someone has to notice the failure, read the logs, bump a config value and start again. In practice people respond by setting caps high enough that they never trigger, which defeats the point.

Option 2: overspend

Let the agent go over "just this once", maybe with a soft limit and an alert.

This keeps the task alive, but now the cap is a suggestion. If the agent can decide to exceed it, so can a buggy plan or an injected instruction. Any limit the agent can talk itself past is not a limit.

Option 3: hand off to a human

The agent stops at the cap, explains what it needs and why, and waits. A human decides. If they approve and pay, the agent continues from where it paused, with its state intact.

This keeps the cap hard and the task alive. It needs three pieces:

  1. A guard that says "ok" or "handoff" before each paid step.
  2. A way to ask a human for exactly the shortfall, with a link they can act on.
  3. A way to resume the paused run safely when the human has acted.

Building option 3

The example repo is send21-over-budget-handoff: Python, MIT, a LangGraph adapter, an offline mock mode and 25 tests. It is an example, not an official LangGraph integration.

1. The guard

The cap check is deliberately boring:

@dataclass
class BudgetGuard:
    cap: Decimal
    spent: Decimal
    currency: str = "USD"

    def check(self, cost: Decimal) -> Literal["ok", "handoff"]:
        return "handoff" if self.spent + cost > self.cap else "ok"

4.20 + 1.50 > 5.00, so the step does not run.

2. The ask

On "handoff", the agent calls the send21 MCP tool create_payment_request with an amount, a payee address, a memo and an orderId. It gets back an id and a payPath, prints the pay link, and pauses with LangGraph's interrupt():

def wait(s: State):
    run = store.get_run(s["run_id"])
    data = interrupt({"pay_link": run["pay_link"], "order_id": run["order_id"]})
    # Resumed by a verified draft.confirmed. Raise the cap by the confirmed amount.
    cap = Decimal(s["cap"]) + Decimal(str(data.get("fiatAmount", run["pending_cost"])))
    return {"cap": str(cap), "handoff_step": None}

The important part is what the agent's key can do. It is created with the drafts:write scope only. send21 prepares payment instructions; the human opens the link and signs in their own wallet, and funds go straight from the payer's wallet to the receiver's wallet. send21 never holds keys or funds, and the agent's key cannot move funds. There is no signing code and no private key in the repo. The only way past the cap is a human signature.

3. The resume

When the payment is confirmed, send21 sends an HMAC-signed draft.confirmed webhook. The receiver checks the signature on the raw body in constant time before parsing anything:

def verify_signature(raw: bytes, header: str | None, secret: str) -> bool:
    if not secret or not isinstance(header, str):
        return False
    return hmac.compare_digest(sign(raw, secret).encode(), header.encode())

Then it stores the delivery id so retries do nothing, matches orderId to the paused run, and resumes it with Command(resume=...). Run status only moves forward, so a late or out-of-order event cannot undo a resume. A draft.amount_mismatch event marks the run for review instead of resuming.

Try it without paying anything

python3.11 -m venv .venv && . .venv/bin/activate
pip install -e ".[dev]"
python scripts/run_demo.py --mock
pytest

Mock mode uses a fake client and a locally signed fake webhook, so it needs no network and no account. The output is the same every run, and it is what the GIF above shows.

Using another framework

The core in handoff/ has no LangGraph imports. To port it, keep the guard, the client and the webhook receiver, and replace the adapter: pause at the handoff, and call resume(run_id, event_data) when the webhook arrives.

What is still open

The README lists what is not yet verified against the live API, for example testnet support for API keys and the exact payPath format. Live mode keeps checkpoints in memory, so use a persistent checkpointer for anything beyond a demo.

Takeaway

A spend cap is only useful if hitting it is a normal, recoverable event. Erroring throws work away, overspending makes the cap meaningless. Handing off keeps the cap hard and puts the decision with a person, who approves with a signature the agent cannot forge.

Repo: https://github.com/send21io/send21-over-budget-handoff
More examples: https://github.com/send21io/send21-examples
send21 MCP: https://send21.io/mcp
Questions: [email protected]

📰 Read the original article on Dev.to AI

Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.