Dev.to Security πŸ” Cybersecurity πŸ‘ 0 πŸ“– 5 min read

What does an AI security review actually cost? I measured it per run

Everyone asks whether AI agents can find security bugs. Fewer people ask the question a team lead asks five minutes later: what does it cost each time we run it? I'm an intern at RakFort in Dublin, and for the last few

Everyone asks whether AI agents can find security bugs. Fewer people ask the question a team lead asks five minutes later: what does it cost each time we run it?

I'm an intern at RakFort in Dublin, and for the last few weeks my job has been testing the cost tracking in secfoo, an open-source (MIT) CLI that gives AI coding agents a fixed security brief and a fixed report format. This post is what I found, including the parts that are not finished.

The problem with "just ask the agent"

You can point Claude Code, Cursor or Codex at a repo and ask for a security review. You will get something useful. You will not get:

  • the same structure twice, so you cannot compare runs
  • a record of what the run cost
  • a way to stop a CI job that is burning money

secfoo handles the first with what it calls skills (structured briefs such as sast, threat-modeling and secret-scanning). The second and third are what I tested.

The setup

I used a deliberately small target: a 30-line Flask app with one SQL injection, where get_user() builds its query by joining strings. Small on purpose, so I could check the report by hand.

pip install "secfoo[api]"
export OPENAI_API_KEY=...
secfoo run --skill sast --agent api --target . --project-name demo --app-id ""

The api agent calls a model API directly, so you do not need an agent CLI installed. The last two flags skip the interactive prompts, which matters in scripts.

Here is the real output:

$ secfoo run --skill sast --agent api --target . --project-name demo --app-id ""
  success β€” SAST β€” Static Code Analysis
                                              Assessment results
┏━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━┳━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┓
┃ Skill                       ┃ Status  ┃ Duration ┃ Findings (C/H/M/L) ┃   Cost ┃ Run ID                               ┃
┑━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━╇━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┩
β”‚ SAST β€” Static Code Analysis β”‚ success β”‚ 11.2s    β”‚ 0/1/0/0            β”‚ <$0.01 β”‚ 888f222e-7228-4b2b-b306-78db7a17d499 β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
AI spend for this run: <$0.01
Run secfoo serve to view full reports in your browser.

What one run cost

The scan finished in 11.2 seconds and reported one finding: the SQL injection in get_user(), which is the right answer for this app.

It used 2,183 input tokens and 981 output tokens, 3,164 in total. secfoo printed the spend for the run as under one cent. It shows <$0.01 rather than rounding down to zero, which I liked: a zero would suggest the run was free.

In my earlier testing, a repeat scan of an unchanged repo used exactly the same input tokens. That is expected today: nothing is cached between runs yet, so the second scan costs the same as the first.

Same kind of target, different agent

Earlier I ran a similar small app through Codex instead of the api agent:

secfoo run --skill sast --agent codex --target . --project-name demo --app-id ""

It finished in 27.5 seconds and also found the injection. Two things were different:

  • Cost shows as -, not $0.00. Codex runs on a subscription login, so there is no per-run price to record. A dash is the right answer. A zero would be a lie.
  • Tokens were much higher. Codex's own log reported 15,201 tokens for the run, roughly five times the api agent, because an agent CLI explores the repo with tools instead of receiving one prepared prompt.

That second point is the most useful thing I learned. The agent you pick can change the token count by a multiple, on a small target, for the same finding.

Seeing spend across runs

Per-run numbers are nice. The summary is what a team would actually use:

secfoo cost
secfoo cost --by skill --since 2026-09-01
secfoo cost --by agent --project demo

You can group by agent, skill or project. The same numbers appear in the local dashboard (secfoo serve), and in my tests the dashboard total matched the CLI total for the same runs.

$ secfoo cost --by agent --project demo
                   AI spend by agent
┏━━━━━━━┳━━━━━━┳━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━┳━━━━━━━━┓
┃ Agent ┃ Runs ┃ Input tokens ┃ Output tokens ┃   Cost ┃
┑━━━━━━━╇━━━━━━╇━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━╇━━━━━━━━┩
β”‚ api   β”‚    2 β”‚        2,183 β”‚           981 β”‚ <$0.01 β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚ Total β”‚    2 β”‚        2,183 β”‚           981 β”‚ <$0.01 β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”˜
1 run(s) have no cost recorded -- their agent doesn't report it (cursor, agy), or they ran before cost tracking existed.

That table shows two runs, not one. My first attempt failed because I pasted my API key with a character missing. secfoo still counted it as a run and told me that one run had no cost recorded. I did not plan that, but it is a fair picture of how it behaves when something goes wrong.

Making cost a CI gate

This is the part I think matters most for real pipelines:

secfoo run --skill sast --agent api --target . \
  --fail-on high --max-cost 2.00 --json > secfoo-result.json

The exit code tells CI what happened: 0 means the run succeeded and no gate tripped, 1 means a run failed or timed out, and 2 means the run succeeded but a finding or the cost crossed your threshold. A security scan that can fail a build for being too expensive is not something I had seen before.

What is not there yet

I was testing, so I was looking for gaps. The honest list:

  • Cost coverage depends on the agent. API-key runs give you tokens and cost. Subscription-based agent CLIs often give you neither, and you see a dash.
  • Codex tokens are not recorded yet. Codex prints its total, but secfoo does not pick it up today. There is an open item for it.
  • Failed runs are counted as runs. They show up in the totals with no cost recorded, as my mistyped key showed.
  • Severity can differ between the CLI and the dashboard. In my run the CLI table counted the finding as High, and the dashboard labelled the same finding Critical. It is a known issue.
  • --since compares against UTC. If you run scans late at night in another timezone, the cut-off may not be where you expect.
  • No savings on repeat scans yet. Reusing results for unchanged code is planned, not shipped.

What I would tell a team

  1. Start with the api agent on one repo, so you get real token and cost numbers from day one.
  2. Run the same repo through the agent your developers already use, and compare the reports side by side.
  3. Put --max-cost in CI before you put the scan on every pull request.

Try it

pip install secfoo
secfoo run --skill sast --agent claude
secfoo serve

The repo is here: https://github.com/secfoo-com/secfoo. If it is useful to you, a star helps a small open-source project a lot, and issues telling us where the output is wrong help even more.

We are also part of the team organising a secfoo hackathon in Dublin in mid-November. If you would like to take part, or to help as a moderator, leave a comment and I will get back to you.

πŸ“° Read the original article on Dev.to Security

Originally published by Dev.to Security. Aggregated on AIWithGhost for educational purposes β€” full credit and traffic to the original publisher.