Dev.to AI 🤖 Ai 👁 0 📖 4 min read

I Audited 5 Agent Frameworks for Cost Visibility. Here's the Scorecard.

Every agent framework promises observability. Traces, spans, dashboards. But ask a harder question — how much did that agent run cost, and who do I bill? — and the answers get thin fast. So I audited five of the most-us

Every agent framework promises observability. Traces, spans, dashboards. But ask a harder question — how much did that agent run cost, and who do I bill? — and the answers get thin fast.

So I audited five of the most-used agent frameworks against one rubric: cost visibility. Not general observability — the specific ability to see, attribute, and control spend. I read the official docs for each (sources linked throughout). No guessing, no marketing pages. Here's the scorecard.

The rubric

Four questions, scored 1–5:

  1. Does it track token usage natively, per run / per agent / per tool call?
  2. Does it convert tokens to dollars?
  3. Can you attribute spend to an agent, session, or user out of the box?
  4. Does it support budgets or cost guardrails natively?

The scorecard

Framework Score Verdict
LangGraph + LangSmith 5/5 The gold standard
Amazon Bedrock AgentCore 5/5 Zero-wiring cost visibility
Strands Agents SDK 4/5 Best guardrails, priced in tokens
CrewAI 3/5 Tokens yes, dollars DIY
AutoGen (v0.4) 2/5 Went backwards

LangGraph + LangSmith: 5/5

LangSmith does the thing everyone else stops short of: it converts tokens to dollars automatically, using a customizable per-model pricing table (docs: https://docs.langchain.com/langsmith/cost-tracking). Custom or non-linear pricing goes through usage_metadata. The trace tree breaks cost down per node, so you can separate what the supervisor spent from what the specialist agents spent. Metadata-based attribution covers users and sessions, and there are dashboards on top.

The blind spot: no native budget enforcement. You can see every dollar with precision — you just can't make the framework stop spending them. Visibility without teeth.

Bedrock AgentCore: 5/5

The managed CloudWatch GenAI dashboard shows per-invocation token usage and cost with zero wiring — no callbacks, no exporters, no pricing tables to maintain. Per-user traceability comes through span attributes, and tenant attribution via OpenTelemetry baggage.

The blind spot: spend control is external. It's CloudWatch alarms plus AWS Budgets — both notify-only. The platform will tell you, beautifully, exactly how much you overspent. It won't pull the plug.

Strands Agents SDK: 4/5

Every AgentResult carries automatic per-invocation token metrics — input, output, total, cache tokens, plus per-tool stats (docs: https://strandsagents.com/docs/user-guide/sdk/observability-evaluation/metrics/). And Strands is the only framework in this group with native budget caps: limits={"turns", "output_tokens", "total_tokens"} halts the agent loop with explicit stop reasons.

The blind spot: the caps are token-denominated, and there's no dollar conversion. A 10,000-token cap means wildly different dollars on Haiku vs. Opus. It's the only framework with a native kill-switch — it just thinks in the wrong currency.

CrewAI: 3/5

token_usage on every kickoff() result is genuinely zero-setup, and per-task usage rides along on TaskOutput. The event bus plus OpenTelemetry export gives you the plumbing to build whatever you want (docs: https://docs.crewai.com/v1.15.12/en/observability/overview).

The blind spot: the framework stops at tokens. No pricing, no dashboard, no budgets. Dollar cost is a DIY project, and per-agent attribution needs custom event listeners. You get the raw material of cost visibility and none of the product.

AutoGen (v0.4): 2/5

Here's the irony: AutoGen had built-in cost logging. Version 0.2 shipped start_logging / print_usage_summary with cost computation. Then came the v0.4 rewrite — and cost visibility went backwards. The current telemetry story is native OpenTelemetry with GenAI semantic-convention spans (docs: https://microsoft.github.io/autogen/stable/user-guide/core-user-guide/framework/telemetry.html): token data exists as span attributes for your backend to aggregate. No first-party cost computation, no dashboard, no budgets.

The blind spot: everything. The most popular multi-agent framework in the ecosystem currently treats cost as somebody else's problem — and it used to be better at this.

Three patterns

1. The industry stops at tokens. Only LangSmith and AgentCore convert to dollars natively. Everyone else hands you token counts and wishes you luck with the pricing page. Tokens are an implementation detail; dollars are the budget. The gap between them is where finance teams lose trust in AI projects.

2. Nobody ships a dollar kill-switch. Strands has token caps. AutoGen has a token-limited context. Everything else is alarms and custom logic. As of today, no major framework lets you say "halt this agent if it spends more than $2" — in dollars, natively, enforced. That missing primitive is the single biggest gap in agent cost governance.

3. Visibility went backwards on at least one framework. AutoGen's regression from v0.2 to v0.4 is a cautionary tale: cost features are the first thing cut in a rewrite because they're nobody's launch-blocking requirement — until the invoice arrives.

What to do about it

If you're picking a framework this quarter: LangSmith or AgentCore if you want cost visibility today; Strands if you want guardrails and can live with token-denominated caps. If you're already on CrewAI or AutoGen, budget a sprint for the DIY layer — token export, a pricing table, per-agent attribution, and an external budget alarm — because the framework isn't going to do it for you.

And framework maintainers: the tokens-to-dollars gap is the most valuable ten lines of code in your roadmap. Ship the pricing table. Ship the dollar cap. Your users' CFOs will thank you.

Methodology: scored against official documentation as of October 2026. Karmendra Pandey is a Practice Architect in AI & ML at TEK Systems, working on cost-governed AI agents on AWS. Reference architecture: https://doi.org/10.5281/zenodo.23178788.

📰 Read the original article on Dev.to AI

Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.