Dev.to AI 🤖 Ai 👁 0 📖 3 min read

The Receipt: I Priced Every Token of One AI Agent Task

Everyone talks about AI costs in aggregate. Dashboards show monthly spend, per-model totals, team budgets. Nobody shows you a receipt. So here's one. A single AI agent task, priced line by line — every token, every tool

Everyone talks about AI costs in aggregate. Dashboards show monthly spend, per-model totals, team budgets. Nobody shows you a receipt.

So here's one. A single AI agent task, priced line by line — every token, every tool call, every retry. This is a worked example from production-style workloads I describe in my reference architecture work (full paper: https://doi.org/10.5281/zenodo.23178788). The numbers are representative of a mid-complexity support-triage agent running on AWS Bedrock. Your numbers will differ. The shape of the receipt won't.

The task

A customer support triage agent: read an incoming ticket, pull the customer's history, check the knowledge base, decide whether to auto-resolve or escalate, and draft the response. One task, one receipt.

The receipt

# Line item Tokens in Tokens out Unit cost basis Cost
1 System prompt (loaded once, cached) 1,850 — cached input $0.0011
2 Ticket text + customer history 2,300 — standard input $0.0069
3 Knowledge base retrieval (3 chunks) 4,100 — standard input $0.0123
4 Reasoning: triage decision — 420 standard output $0.0050
5 Tool call: CRM lookup 180 90 in/out $0.0016
6 Reasoning: draft response — 680 standard output $0.0082
7 Retry: first draft failed tone eval, regenerated 6,200 710 in (cached) + out $0.0121
8 Embedding: ticket for memory store 340 — embedding $0.0000
9 Final response + audit log write — 350 standard output $0.0042
Total ~15,000 ~2,250 ~$0.051

Wait — $0.051, not $0.23? The $0.23 figure from my paper is the fully-loaded cost: it includes the amortized cost of the eval harness, the retrieval infrastructure, and the human review queue time for escalations. The raw inference receipt is five cents. The governed cost is twenty-three cents. Both numbers matter, and most teams track neither.

Reading the receipt

Three things jump out:

1. The retry cost more than the original attempt. Line 7 — a failed tone evaluation forced a regeneration — cost $0.0121, more than lines 4+6 combined. Retries are the silent budget killer in agent systems. Every eval gate you add has a cost; every gate you skip has a bigger one. The receipt lets you price the tradeoff instead of guessing.

2. Context is the bulk of input spend. Lines 2+3 are 6,400 input tokens — 43% of the total input. That knowledge-base retrieval is doing real work, but it's also the first place to optimize: better chunking, semantic caching of frequent queries, and prompt compression all attack this line directly.

3. The system prompt is nearly free — because it's cached. Line 1 costs a tenth of a cent thanks to prompt caching. Without caching, that 1,850-token system prompt gets re-priced on every call in the chain. If your provider supports caching and you're not using it, you're donating money.

Why the receipt is the unit of cost governance

You can't govern what you can't itemize. Monthly dashboards tell you that you spent $40K. The receipt tells you why — which line items to attack, which eval gates earn their keep, which retries are worth preventing.

Three practices that fall out of this:

  • Emit a receipt per task. Log input/output tokens per span — per agent, per tool call, per retry. If your framework doesn't do this natively, add a callback handler.
  • Budget at the task level, not the month level. A per-task budget with a circuit breaker ("halt if projected cost exceeds $X") catches runaway loops in seconds instead of at invoice time.
  • Price your eval gates. Every quality check costs tokens. The receipt tells you whether the gate is cheaper than the failure it prevents.

The next time someone asks what your AI agents cost, don't open the dashboard. Show them a receipt.

Karmendra Pandey is a Practice Architect in AI & ML at TEK Systems, building production agentic AI on AWS. His reference architecture for cost-governed AI agents is at https://doi.org/10.5281/zenodo.23178788, and the companion implementation at https://github.com/karmendra8386/agentic-ai-aws-reference.

📰 Read the original article on Dev.to AI

Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.