Dev.to AI ๐Ÿค– Ai ๐Ÿ‘ 0 ๐Ÿ“– 5 min read

Security Benchmark Explorer: An AI Agent That Only Works Because the Content Is Structured

This is a submission for the Sanity Challenge, Path One: Ship an Agent That Queries Real Content What I Built Security Benchmark Explorer is an AI agent that helps security engineers and AI researchers explor

This is a submission for the Sanity Challenge, Path One: Ship an Agent That Queries Real Content

What I Built

Security Benchmark Explorer is an AI agent that helps security engineers and AI researchers explore a 54-case benchmark of vulnerability detection test cases โ€” including 5 deliberate false-positive traps designed to distinguish data-flow reasoning from pattern matching.

A generic LLM knows what SQL injection is. It does not know which of my 54 cases are vulnerable, which are safe, what difficulty tier each falls into, or why a specific case was classified the way it was. That knowledge lives in a structured Sanity dataset, and the agent reads it through Sanity Context MCP.

The agent answers five kinds of questions that a keyword search cannot:

  • Exact case lookup: "Explain the rationale for EVAL-004" โ†’ requires a caseId == "EVAL-004" filter
  • Compound filtering: "Show me all hard-difficulty SQL Injection cases" โ†’ requires category == "SQL Injection" && difficulty == "hard"
  • Pattern matching on IDs: "Which cases are false-positive traps?" โ†’ requires caseId match "TRAP-*"
  • Field-by-field comparison: "What's the difference between CMD-004 and CMD-005?" โ†’ requires reading two documents and comparing the vulnerable and rationale fields
  • Aggregation over structured data: "How many vulnerable cases are in the benchmark?" โ†’ requires count() on a boolean field

Try it: security-benchmark-explorer.vercel.app/explorer

Demo

Deployed app: security-benchmark-explorer.vercel.app/explorer

Query 1 โ€” Exact case lookup

You: Explain the rationale for EVAL-004

Agent: EVAL-004 is a hard-difficulty Code Injection case. The code is expression = OPERATIONS[user_input] followed by result = eval(expression). It's vulnerable because the attacker controls the lookup key and therefore potentially selects executable expressions.

A keyword search for "EVAL-004" returns nothing โ€” the case ID only exists in a structured field.

Query 2 โ€” Compound filter across two fields

You: Show me all hard-difficulty SQL Injection cases

Agent: The hard-difficulty SQL Injection cases are SQL-004 and SQL-005. SQL-004 is vulnerable because it parameterizes the query values but still interpolates the column name from user input. SQL-005 uses the same structure but validates the column against an allowlist, making it safe.

This is a GROQ query: *[_type == "benchmarkCase" && category == "SQL Injection" && difficulty == "hard"]. Two field filters, one result set. A keyword search for "hard SQL Injection" returns nothing useful โ€” it can't distinguish the vulnerable case from the safe one in the same category.

Query 3 โ€” The false-positive traps

You: Which cases are false-positive traps?

Agent: The 5 false-positive traps are TRAP-001 through TRAP-005 โ€” all safe code designed to look dangerous. TRAP-001 is subprocess.run(["ls", user_input]) โ€” safe because no shell is invoked. TRAP-002 uses secrets.token_hex(32). TRAP-003 parameterizes both values. TRAP-004 evals a constant. TRAP-005 uses MD5 for a file checksum, not password hashing.

A keyword search for "false positive" returns zero results โ€” the phrase doesn't appear in any document text.

Query 4 โ€” Two-case comparison

You: What's the difference between CMD-004 and CMD-005?

Agent: CMD-004 is subprocess.run(["cat", filename], check=True) โ€” safe, because user input is passed as an argument, not shell syntax. CMD-005 is command = ["cat", user_input]; subprocess.run(command, shell=True) โ€” vulnerable, because shell=True changes the security properties of the invocation. The only difference is shell=True.

This requires reading two documents and comparing their vulnerable and rationale fields.

Query 5 โ€” Aggregation

You: How many vulnerable cases are in the benchmark?

Agent: 26 of the 54 benchmark cases are vulnerable. 28 are safe.

This is count(*[_type == "benchmarkCase" && vulnerable == true]). A Knowledge Base of prose documents can't answer this. Only a structured dataset can.

Code

Repository: github.com/CrossDomainSolutionArchitect/kaggle-security-benchmark-explorer

The repo contains the Sanity Studio (in studio/), the Next.js agent app (in web/), and the 54-case dataset in NDJSON format (in data/).

Key files:

  • studio/schemaTypes/benchmarkCase.ts โ€” the schema (6 fields)
  • data/benchmark_cases.ndjson โ€” the 54-case dataset
  • web/src/lib/sanity-context.ts โ€” the MCP client
  • web/src/app/api/explorer/route.ts โ€” the agent route
  • web/src/app/explorer/page.tsx โ€” the chat UI

How I Used Sanity

I built one document type โ€” benchmarkCase โ€” with six fields:

Field Type Purpose
caseId string e.g., SQL-001, CMD-004, TRAP-002
code text The Python snippet
vulnerable boolean Ground-truth label
category string Vulnerability type
difficulty string easy / medium / hard
rationale text Why this classification

The vulnerable boolean is the critical field. Without it, you can't distinguish subprocess.run(["cat", filename]) (safe) from subprocess.run("cat " + filename, shell=True) (vulnerable). Both look similar. Only the structured ground-truth label separates them.

I pointed Sanity Context at the benchmarkCase dataset source, not a Knowledge Base. This was a deliberate choice: the challenge's use case was structured lookups, and Knowledge Bases serve semantic search over prose. My data has exact field values (caseId, category, difficulty, vulnerable) that require exact-match filtering, not similarity search.

The Sanity Context MCP endpoint exposes the GROQ tools to the agent. The agent calls groq_query with a filter, gets structured documents back, and reasons over them. It never guesses โ€” it always queries.

The endpoint's custom instructions tell the agent to always cite case IDs, reference the 5 TRAP cases for false-positive questions, reference CMD-004 and TRAP-001 for the variable-name bias finding, and never invent case IDs.

Sanity Project Details

  • Project ID: 2el69ou3
  • Organization ID: o9cl071wg
  • Dataset: production (public)
  • Document type: benchmarkCase
  • Total documents: 54
  • MCP endpoint: security-benchmark-mcp

Public dataset URL:

View all 54 cases in the Sanity Content Lake

You can also view the deployed Studio at:
kaggle-security-benchmark-explorer.sanity.studio

Agent Session

I built this agent in GitHub Codespaces without using an AI coding CLI, so I have no Agent Session transcript to embed. The full build history is visible in the repository's commit log.

Why This Only Works Because the Content Is Structured

This is the challenge's core requirement, so let me be explicit about it.

  • A keyword search for "EVAL-004" returns nothing. The case ID only exists in a structured field.
  • A keyword search for "hard SQL Injection" returns nothing useful. It can't filter on difficulty == "hard" and category == "SQL Injection" simultaneously.
  • A keyword search for "false positive" returns nothing. The phrase doesn't appear in any document text โ€” the TRAP cases are identified by their caseId pattern.
  • A question like "How many vulnerable cases are there?" can only be answered by counting a boolean field. No amount of prose can support that.

The agent works because the data is structured. A keyword search over the same content would return nothing.

What I'd Build Next

Three additions I'd make with more time:

  1. Model performance as a related document type. Each case could reference the models that failed it. This would let the agent answer "Which models failed TRAP-001?" โ€” a query that spans two structured document types.

  2. Vulnerability relationships as references. Model relatedVulnerabilities between cases so the agent can answer "What other cases are similar to SQL-004?" using graph traversal.

  3. App SDK interface. Build a Studio app that shows the benchmark as a filterable table, so security teams can explore the data without the chat interface.

Built With

Sanity Context MCP ยท Next.js 16 ยท Vercel AI SDK ยท Gemini 2.5 Flash ยท TypeScript ยท Tailwind CSS

Built by

vernard_sharbney_4c39f22b

sanitychallenge

๐Ÿ“ฐ Read the original article on Dev.to AI

Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes โ€” full credit and traffic to the original publisher.