Security Benchmark Explorer: An AI Agent That Only Works Because the Content Is Structured
This is a submission for the Sanity Challenge, Path One: Ship an Agent That Queries Real Content What I Built Security Benchmark Explorer is an AI agent that helps security engineers and AI researchers explor
This is a submission for the Sanity Challenge, Path One: Ship an Agent That Queries Real Content
What I Built
Security Benchmark Explorer is an AI agent that helps security engineers and AI researchers explore a 54-case benchmark of vulnerability detection test cases โ including 5 deliberate false-positive traps designed to distinguish data-flow reasoning from pattern matching.
A generic LLM knows what SQL injection is. It does not know which of my 54 cases are vulnerable, which are safe, what difficulty tier each falls into, or why a specific case was classified the way it was. That knowledge lives in a structured Sanity dataset, and the agent reads it through Sanity Context MCP.
The agent answers five kinds of questions that a keyword search cannot:
-
Exact case lookup: "Explain the rationale for EVAL-004" โ requires a
caseId == "EVAL-004"filter -
Compound filtering: "Show me all hard-difficulty SQL Injection cases" โ requires
category == "SQL Injection" && difficulty == "hard" -
Pattern matching on IDs: "Which cases are false-positive traps?" โ requires
caseId match "TRAP-*" -
Field-by-field comparison: "What's the difference between CMD-004 and CMD-005?" โ requires reading two documents and comparing the
vulnerableandrationalefields -
Aggregation over structured data: "How many vulnerable cases are in the benchmark?" โ requires
count()on a boolean field
Try it: security-benchmark-explorer.vercel.app/explorer
Demo
Deployed app: security-benchmark-explorer.vercel.app/explorer
Query 1 โ Exact case lookup
You: Explain the rationale for EVAL-004
Agent: EVAL-004 is a hard-difficulty Code Injection case. The code is
expression = OPERATIONS[user_input]followed byresult = eval(expression). It's vulnerable because the attacker controls the lookup key and therefore potentially selects executable expressions.
A keyword search for "EVAL-004" returns nothing โ the case ID only exists in a structured field.
Query 2 โ Compound filter across two fields
You: Show me all hard-difficulty SQL Injection cases
Agent: The hard-difficulty SQL Injection cases are SQL-004 and SQL-005. SQL-004 is vulnerable because it parameterizes the query values but still interpolates the column name from user input. SQL-005 uses the same structure but validates the column against an allowlist, making it safe.
This is a GROQ query: *[_type == "benchmarkCase" && category == "SQL Injection" && difficulty == "hard"]. Two field filters, one result set. A keyword search for "hard SQL Injection" returns nothing useful โ it can't distinguish the vulnerable case from the safe one in the same category.
Query 3 โ The false-positive traps
You: Which cases are false-positive traps?
Agent: The 5 false-positive traps are TRAP-001 through TRAP-005 โ all safe code designed to look dangerous. TRAP-001 is
subprocess.run(["ls", user_input])โ safe because no shell is invoked. TRAP-002 usessecrets.token_hex(32). TRAP-003 parameterizes both values. TRAP-004 evals a constant. TRAP-005 uses MD5 for a file checksum, not password hashing.
A keyword search for "false positive" returns zero results โ the phrase doesn't appear in any document text.
Query 4 โ Two-case comparison
You: What's the difference between CMD-004 and CMD-005?
Agent: CMD-004 is
subprocess.run(["cat", filename], check=True)โ safe, because user input is passed as an argument, not shell syntax. CMD-005 iscommand = ["cat", user_input]; subprocess.run(command, shell=True)โ vulnerable, becauseshell=Truechanges the security properties of the invocation. The only difference isshell=True.
This requires reading two documents and comparing their vulnerable and rationale fields.
Query 5 โ Aggregation
You: How many vulnerable cases are in the benchmark?
Agent: 26 of the 54 benchmark cases are vulnerable. 28 are safe.
This is count(*[_type == "benchmarkCase" && vulnerable == true]). A Knowledge Base of prose documents can't answer this. Only a structured dataset can.
Code
Repository: github.com/CrossDomainSolutionArchitect/kaggle-security-benchmark-explorer
The repo contains the Sanity Studio (in studio/), the Next.js agent app (in web/), and the 54-case dataset in NDJSON format (in data/).
Key files:
-
studio/schemaTypes/benchmarkCase.tsโ the schema (6 fields) -
data/benchmark_cases.ndjsonโ the 54-case dataset -
web/src/lib/sanity-context.tsโ the MCP client -
web/src/app/api/explorer/route.tsโ the agent route -
web/src/app/explorer/page.tsxโ the chat UI
How I Used Sanity
I built one document type โ benchmarkCase โ with six fields:
| Field | Type | Purpose |
|---|---|---|
caseId |
string | e.g., SQL-001, CMD-004, TRAP-002
|
code |
text | The Python snippet |
vulnerable |
boolean | Ground-truth label |
category |
string | Vulnerability type |
difficulty |
string | easy / medium / hard |
rationale |
text | Why this classification |
The vulnerable boolean is the critical field. Without it, you can't distinguish subprocess.run(["cat", filename]) (safe) from subprocess.run("cat " + filename, shell=True) (vulnerable). Both look similar. Only the structured ground-truth label separates them.
I pointed Sanity Context at the benchmarkCase dataset source, not a Knowledge Base. This was a deliberate choice: the challenge's use case was structured lookups, and Knowledge Bases serve semantic search over prose. My data has exact field values (caseId, category, difficulty, vulnerable) that require exact-match filtering, not similarity search.
The Sanity Context MCP endpoint exposes the GROQ tools to the agent. The agent calls groq_query with a filter, gets structured documents back, and reasons over them. It never guesses โ it always queries.
The endpoint's custom instructions tell the agent to always cite case IDs, reference the 5 TRAP cases for false-positive questions, reference CMD-004 and TRAP-001 for the variable-name bias finding, and never invent case IDs.
Sanity Project Details
-
Project ID:
2el69ou3 -
Organization ID:
o9cl071wg -
Dataset:
production(public) -
Document type:
benchmarkCase - Total documents: 54
-
MCP endpoint:
security-benchmark-mcp
Public dataset URL:
View all 54 cases in the Sanity Content Lake
You can also view the deployed Studio at:
kaggle-security-benchmark-explorer.sanity.studio
Agent Session
I built this agent in GitHub Codespaces without using an AI coding CLI, so I have no Agent Session transcript to embed. The full build history is visible in the repository's commit log.
Why This Only Works Because the Content Is Structured
This is the challenge's core requirement, so let me be explicit about it.
- A keyword search for "EVAL-004" returns nothing. The case ID only exists in a structured field.
- A keyword search for "hard SQL Injection" returns nothing useful. It can't filter on
difficulty == "hard"andcategory == "SQL Injection"simultaneously. - A keyword search for "false positive" returns nothing. The phrase doesn't appear in any document text โ the TRAP cases are identified by their
caseIdpattern. - A question like "How many vulnerable cases are there?" can only be answered by counting a boolean field. No amount of prose can support that.
The agent works because the data is structured. A keyword search over the same content would return nothing.
What I'd Build Next
Three additions I'd make with more time:
Model performance as a related document type. Each case could reference the models that failed it. This would let the agent answer "Which models failed TRAP-001?" โ a query that spans two structured document types.
Vulnerability relationships as references. Model
relatedVulnerabilitiesbetween cases so the agent can answer "What other cases are similar to SQL-004?" using graph traversal.App SDK interface. Build a Studio app that shows the benchmark as a filterable table, so security teams can explore the data without the chat interface.
Built With
Sanity Context MCP ยท Next.js 16 ยท Vercel AI SDK ยท Gemini 2.5 Flash ยท TypeScript ยท Tailwind CSS
Built by
vernard_sharbney_4c39f22b
sanitychallenge
Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes โ full credit and traffic to the original publisher.