Dev.to Security πŸ” Cybersecurity πŸ‘ 0 πŸ“– 4 min read

Monitoring AI coding agent costs and activity on your own machine

I run Claude Code, Cursor, Codex and a handful of sub-agents on my own machines every day. For a long time, I knew what they cost but had little visibility into their work. If an agent felt slow or got stuck, I had to d

Monitoring AI coding agent costs and activity on your own machine

I run Claude Code, Cursor, Codex and a handful of sub-agents on my own machines every day.

For a long time, I knew what they cost but had little visibility into their work. If an agent felt slow or got stuck, I had to dig through separate logs to understand why.

That became harder as I started using several agents on the same project. Each kept its own session history, and there was no single place to review the activity.

We built AgentMon to monitor agents across a company. AgentMon Start brings that visibility to a desktop app for individual developers and teams.

AgentMon Start Dashboard

I’m part of the team behind AgentMon. This post explains how Start collects data and what you can inspect after installation.

How collection works

AgentMon Start reads the session logs your coding agents already write. It supports more than 16 agents, harnesses and LLMs, including Claude Code, Claude Desktop, Codex, GitHub Copilot, Cursor, Gemini CLI and Goose.

There is no SDK to add to your project or proxy to configure.

The app has three local components:

  • A collector that parses agent logs into snapshots of token usage, costs, tool calls, errors and security findings.
  • A local server that stores those snapshots in SQLite and analyzes them.
  • A desktop dashboard for viewing the results.

The local server listens on 127.0.0.1. Closing the dashboard leaves collection running in the background.

Data collection

Start with costs

The Overview shows spending alongside session counts, token usage and cache hit rate. You can inspect the last 24 hours, seven days, 30 days or all recorded activity.

The useful question is usually more specific than β€œHow much did I spend?”

Which session used the budget? Which model was running? Did an agent spend time retrying the same operation?

The dashboard lets you move from the total to the sessions behind it. The hourly chart helps you locate a burst of activity, while the model breakdown shows how spending was distributed.

Inspect the sessions and tool calls

Sessions groups activity into threads representing continuous work on a repository.

AgentMon Start Session view

Each thread shows the agent, cost, token usage and activity mix, including file reads, edits, shell commands and web access. Errors and security findings appear alongside the session.

In the example from our launch post, a Claude Code thread cost $1.96 and ran 26 Bash calls. Two calls failed, and the thread was flagged for a prompt-injection finding.

Those details give you somewhere to start investigating. You can review the failed calls and decide whether the agent needs another instruction, a different tool or a stopped session.

The Tools view also aggregates failures across sessions, so you can see whether one tool accounts for most of the errors.

AgentMon Start Agents Tool Use

Review a trace

When an agent stalls or produces an unexpected result, open its session trace.

The waterfall shows recorded LLM calls, tool calls and reasoning steps with timings. The side panel provides the available input, output, metadata and raw JSON.

This helps you follow the recorded sequence of events and inspect where the session went wrong.

AgentMon Start Trace waterfall

Set alerts and review recommendations

Alerts check every 60 seconds. Depending on your plan, they can flag cost thresholds, spending spikes, error spikes, long-running sessions and unmanaged agent activity.

Recommendations examine the last 30 days for cost and reliability problems, including low cache use, repeated retries, wasted sessions and bloated prompts. Each recommendation includes evidence and an estimated saving.

Security checks flag secrets in prompts, prompt-injection attempts and dangerous shell commands. A finding gives you something to investigate; review the session context before deciding what action to take.

Local storage and optional team sharing

Activity data stays on your machine by default.

The app contacts Codenotary for sign-in, license checks and feedback you choose to send. Connecting a device to a team dashboard is an explicit step, and the app shows what will be uploaded before syncing begins.

For teams, that shared view helps answer questions about spending across people, projects and models, as well as which agents are running.

Try it on your own sessions

AgentMon Start is available for Windows, macOS and Linux. Your first three devices are free for six months, and paid plans start at $15 per user per month.

Install the app, sign in and continue working. The collector picks up existing agent sessions, and the dashboard fills in from that activity.

A useful first check is to open your most expensive session and look at its tool calls. Was the agent making progress, waiting on something or repeating failed work?

Read the full walkthrough and get started with AgentMon Start.

If you use several coding agents, how do you currently track their costs and investigate failed sessions? I’d be interested to hear what information you’re missing.

πŸ“° Read the original article on Dev.to Security

Originally published by Dev.to Security. Aggregated on AIWithGhost for educational purposes β€” full credit and traffic to the original publisher.