Should This AI Agent Go Live? A 10-Question Go/No-Go Check
Most AI agent reviews end with a vague "looks fine." Then the agent gets a few more tools, a broader token, a memory store, and nobody re-checks. The problem isn't a lack of frameworks. OWASP published the Top 10 for Ag
Most AI agent reviews end with a vague "looks fine." Then the agent gets a few more tools, a broader token, a memory store, and nobody re-checks.
The problem isn't a lack of frameworks. OWASP published the Top 10 for Agentic Applications (2026), NIST has the AI RMF, and every vendor has a whitepaper. The problem is turning those into a decision: should this specific agent go live, and under what conditions?
Here's the lightweight process I use. It fits on one page and you can run it this week.
Step 1: Classify autonomy before you assess anything
The same control gap means very different things for a chatbot and for an agent that can send payments. So start by placing each agent on a simple autonomy ladder:
| Level | What the agent can do | Example |
|---|---|---|
| L0 | Answers only, no tools | FAQ assistant |
| L1 | Reads data through tools | RAG over internal docs |
| L2 | Proposes actions, human executes | Drafts emails, tickets |
| L3 | Executes actions with approval | Sends email after a click |
| L4 | Executes actions autonomously | Auto-remediation, scheduled agents |
Rule of thumb: every level up doubles the evidence you should demand. An L0 bot can ship with basic guardrails. An L4 agent needs proven identity, approval, logging and a kill switch.
Step 2: Inventory what the agent can actually touch
For each agent, write down four things:
- Tools / MCP servers it can call (and who approved each one)
- Identity it acts with: the end user's, or a shared service account?
- Data it can read and write, including memory and vector stores
- Triggers: does a human start it, or can an email, webhook or schedule start it?
If nobody can answer one of these in under five minutes, that's your first finding.
Step 3: Ask for evidence, not opinions
Most reviews fail here. "We have a guardrail" is an opinion. A screenshot of the guardrail blocking a planted canary is evidence. These ten questions cover the controls I see fail most often:
-
Untrusted content: If a document, web page or tool output contains instructions, does the agent ignore them? Evidence: a planted
CANARYinstruction that did not execute. - Tool approval: Are tool descriptions and launch configs pinned, so a changed MCP server is re-approved? Evidence: a modified config that triggered a re-prompt.
- Least privilege: Does each tool run with the minimum scope, not an admin token? Evidence: token scopes listed per tool.
- User identity: Is authorization enforced at the resource with the end user's identity? Evidence: a low-privilege test user denied an admin record.
- Human approval: Do write, delete and send actions require approval that shows the raw tool name and arguments? Evidence: approval screen capture.
- Memory hygiene: Can one session poison memory that another session later trusts? Evidence: cross-session canary test.
- Output handling: Is model output escaped and validated before it reaches a UI, shell or API? Evidence: an HTML/JS canary rendered as text.
- Logging: Is every tool call logged with user ID, arguments and result, outside the agent's own control? Evidence: one reconstructed session from logs.
- Limits: Are there token, cost, rate and step limits per user and per run? Evidence: a capped run and the alert it fired.
- Kill switch: Can you stop the agent and revoke its credentials within your target time? Evidence: a timed drill.
Score each one PASS / PARTIAL / FAIL. Partial means the control exists but you couldn't prove it.
Step 4: Turn results into a decision
A percentage score hides what matters. Use a decision table instead:
| Decision | When |
|---|---|
| β Deploy | All ten PASS for L3βL4; no FAIL on Q1, Q4, Q5, Q10 for L1βL2 |
| β οΈ Deploy with conditions | Only PARTIALs, each with an owner and a date |
| π§ͺ Pilot only | One FAIL outside the critical four, limited users, extra monitoring |
| β Do not deploy | Any FAIL on untrusted content, identity, approval or kill switch at L3+ |
This is the part leadership actually reads. "Agent B: deploy with conditions, 2 items due in 30 days" is a sentence a CISO can sign.
Step 5: Re-run on every change
New tool, new model, new data source, new MCP server version: re-run the ten questions. Keep stable test IDs so results are comparable over time. A control you haven't re-tested since the last change is a hope, not a control.
Common mistakes I keep seeing
- Reviewing the prompt instead of the permissions. Most real agent incidents are authorization bugs, not clever prompts.
- One test account. You can't find identity or cross-tenant failures with a single user.
- No owner per finding. A finding without a name and a date doesn't get fixed.
- Assessing once at launch. Agents drift as tools and data get added.
Want the full version?
The ten questions above are the core. I packaged the complete process as the Agentic AI Security & Governance Toolkit 2026:
- Excel workbook with 72 controls across 14 domains, mapped to OWASP Agentic Top 10 (ASI01βASI10) and NIST AI RMF, assessing up to 5 agents side by side with a deployment decision per agent
- 40 red-team scenarios with a testing playbook and the canary technique
- MCP security review guide and 18 vendor / MCP due-diligence questions
- Governance policy, incident response runbooks and an executive report template, plus a worked example
There's also a Consultant edition with a proposal/SOW, rules of engagement, a pricing calculator and white-label reports if you assess agents for clients.
Whether you use it or not, try the ten questions on one agent you already run. You'll learn more in an hour than from another framework PDF.
Which of the ten would your agent fail today? I'm curious which one comes up most.
For authorized defensive testing only. Test systems you own or have written permission to assess. Not affiliated with OWASP or NIST. Disclosure: I'm the author of the toolkit linked above.
Originally published by Dev.to Security. Aggregated on AIWithGhost for educational purposes β full credit and traffic to the original publisher.