Should This AI Agent Be Allowed to Pay? Designing the Approval Layer Between Agents and Company Money
Disclosure: I'm part of the PinkWallet team building this. Your purchasing agent knows the coffee beans run out Thursday. Your engineering agent knows the API credits are gone. Your ads agent knows a creator batch is re
Disclosure: I'm part of the PinkWallet team building this.
Your purchasing agent knows the coffee beans run out Thursday. Your engineering agent knows the API credits are gone. Your ads agent knows a creator batch is returning 4.6x and wants to put more budget behind it. Each of them stops at the same wall: someone has to pay, and right now that someone is a human clicking a button.
There are two common ways teams handle this today, and both are bad.
Give the agent a company card. Now it can spend on anything, and so can a bug, a leaked key, or a bad prompt. Route every payment to a person. The whole point of an agent acting without waiting for you disappears, and finance ends up approving $40 purchases all day.
This post is a design walkthrough of a third option: an approval layer that sits between the agent and the money, and decides β per request β whether the agent may pay. It's the design behind a product I work on, Pink Agentic AI Payments (by PinkWallet), but the decision logic below is useful whether or not you ever touch our product. I'll point out where a claim comes from our own walkthrough material so you can tell design thinking from vendor pitch.
The one question
Every payment request an agent makes reduces to one question: may this agent pay this? There are exactly three answers:
- ALLOW β the policy already covers this, no human needed
- ASK A PERSON β a rule says this needs a human, and names who
- BLOCK β nothing permits it, so it doesn't happen
Notice there's no fourth answer. A request is never left pending with nobody responsible for it, and it never defaults to yes because nobody got around to saying no.
The decision path
Before any rule fires, a request has to clear a sequence of gates. In the product walkthrough this is described as: registered β active β within budget β within the daily ceiling β vault funded β then each rule, top to bottom, first match wins.
Walking through that in order:
- Registered β is this a known agent with a key, or an unregistered caller? Unregistered gets nothing.
- Active β has this agent been paused? A paused agent is blocked on its very next request β useful when, say, a company pauses its recruiting agent during a hiring freeze.
- Within budget β has the agent's monthly budget already been spent?
- Within the daily ceiling β is there a per-day cap, and has it been hit?
- Vault funded β is there money in the vault this agent is attached to? Agents can only spend from the vault they're attached to; a reserve vault with no agents attached is money no agent can ever reach.
- Rules, top to bottom, first match wins β fraud checks first (bank-detail changes, duplicate invoices), then amount tiers, then domain-specific rules.
- Anything uncovered is blocked. No matching rule means no payment.
Here's a small illustrative policy showing what a few of those rules might look like, built only from example rules in the product walkthrough's sample data (a fictional 50-person startup). This is not our API β it's a pseudo-policy to show the shape of the logic:
# illustrative pseudo-policy, not our API
agent: engineering-ci-bot
vault: engineering-opex
daily_ceiling_usd: 200
rules:
- match: { payee_type: "known_vendor", amount_lt: 1000 }
decision: ALLOW
- match: { amount_between: [1000, 5000] }
decision: ASK
approver: cfo
- match: { amount_gt: 5000 }
decision: ASK
approver_group: "2_of_3_leadership"
- match: { category: "ad_spend", monthly_total_gt: 20000 }
decision: BLOCK
default: BLOCK # anything not matched above
Two defaults are doing the real safety work here, and neither is exotic engineering β they're just refusing to be clever:
Uncovered is blocked, not allowed. If you write 30 rules and a request doesn't match any of them, the safe failure mode is "nothing happens," not "probably fine." An agent can encounter a payee, category, or amount nobody anticipated; the system shouldn't improvise on your behalf.
No answer is a no. When a rule routes a request to a human and the approval window times out, the request is declined, not auto-approved. Silence never spends money. This matters specifically because agents can act at 3 a.m. when nobody's watching a queue.
Rules that need evidence, not just a number
A dollar threshold alone is a weak control β it tells you nothing about whether the purchase is legitimate. The rule set in the walkthrough attaches conditions beyond amount: a matching purchase order, a signed brief, a warehouse scan confirming a return actually happened, a statement match, the currency, the time of day.
The sharpest example is a bank-detail change. If a payee's bank details change recently, that request gets flagged and routed to a specific approver rather than cleared automatically β because a routine-looking invoice with new payee bank details is exactly what invoice-redirection fraud looks like. The rule isn't "is this a known supplier," it's "is this a known supplier paid the way we've always paid them." That distinction is the whole control.
Approvals that don't drown finance
If every exception pages someone, you've rebuilt "route everything to a person" with extra steps. A few design choices keep the human side light:
- Only the exceptions reach a person β not every payment, just the ones a rule routed to ASK, along with the rule that stopped it, the agent's own stated reason, and whatever evidence it attached.
- Approval groups, like 2 of 3 executives or 1 of 2 finance on-call, for amounts big enough to want more than one signature.
- You can only approve what's yours β the policy applies to people too; an approver not named on a request can't act on it, only nudge the person who can.
- A phone that does one thing. The mobile approval flow is deliberately limited to approve/decline β no balances, no transfers, no rule editing. A stolen phone can approve or decline something already asked of it and nothing more. Approving issues the credential immediately; declining logs the approver's name and stops the agent.
Why the credential is a card or a bank transfer
When a request is approved, what gets issued to the agent is a normal payment credential: a single-use virtual card locked to that specific payee and amount, or a bank transfer. Not a crypto wallet, not a new merchant integration the payee has to set up. The reasoning, in the walkthrough's own words: "Fiat first... No merchant integration is needed: what the agent receives is a normal card or a normal bank transfer... No protocol has to win first."
That's a deliberate scoping choice, and it's worth being precise about what it does and doesn't cover. Protocols like AP2, x402, and ACP define how an agent transmits a payment instruction β the wire format and handshake. They don't define who inside a company is allowed to let an agent spend in the first place. Those are different layers of the same problem, and a governance layer can sit above any of them. We wrote a longer comparison of the three protocols if you want the wire-level detail: x402 vs AP2 vs ACP.
If you're specifically working out per-agent budgets or evaluating spend-governance tools, two more of our writeups cover that ground directly: how to set per-agent spending limits and a comparison of AI agent spend-governance tools, plus a dataset we maintain: agent-spending-controls-crosswalk.
The bug case
Here's a scenario from one of the prototype's sample companies (fictional test data, not a real customer) that shows why the daily ceiling matters as much as the monthly budget. At 3:14 a.m., a retry loop in a batch job kept re-firing a request to buy API credits, racking up an attempted $49,600. The agent's $200-a-day rule blocked it β with no approval path, since a bug at 3 a.m. isn't a judgment call, it's a ceiling. The agent was frozen and its owner got a call. Nobody was awake, and nothing was lost.
That's the case for a hard daily ceiling sitting underneath the rule engine, separate from any single rule: a runaway loop shouldn't need a human to notice the pattern before it gets stopped.
Try it
The design above is implemented as a working front-end prototype with a rule engine behind it (the Policy Copilot, the "test a payment" tool, and the day-simulation view all share the same engine). It's public and you can click through it yourself:
Switch companies (top right) to see the coffee shop, the 50-person startup, or the 120-person e-commerce company's rule sets. Press "Simulate a day" to watch requests get allowed, asked, or blocked in sequence. Open Policy Copilot to see how a plain-language sentence like "Ops AI can pay contractors up to $3,000, above that the CFO approves" becomes rules.
To be plain about where this stands: the prototype above shows the console pages, the mobile approval flow and the rule engine with three sample companies. Since 2026-09-30 there is also a public sandbox at agentic-sandbox.pinkwallet.com: a free workspace with an MCP server (7 tools) and a REST API, where you can connect your own agent and watch each payment come back allowed, sent for approval or blocked. Sandbox keys are test credentials and no money moves; production (real single-use cards and bank transfers) is not available yet, and there are no production customers. Connection guides: pinkwallet.com/agentic/developers.
Pink Agentic AI Payments (by PinkWallet) is the approval layer between AI agents and company money: plain-language rules, per-agent budgets, and human approvals decide each payment before a single-use card or bank transfer is issued.
Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes β full credit and traffic to the original publisher.