Dev.to Security 🔐 Cybersecurity 👁 0 📖 4 min read

AgentGuardBench: A Multilingual Security Benchmark for Responsible AI Agents

AI agents are moving beyond generating text. They can retrieve information, retain context, call tools and initiate actions. That makes them useful—but it also changes the nature of AI risk. A harmful output may no long

AI agents are moving beyond generating text. They can retrieve information, retain context, call tools and initiate actions. That makes them useful—but it also changes the nature of AI risk.

A harmful output may no longer remain a sentence on a screen. It could become an unauthorised email, an improper disclosure of personal data, a destructive tool call or an action taken without human approval.

To help teams evaluate these risks in a transparent and reproducible way, I created AgentGuardBench, an open multilingual benchmark for privacy, security and responsible behaviour in tool-using AI agents.

What AgentGuardBench evaluates

AgentGuardBench tests whether an agent can:

  • preserve instruction priority when untrusted content contains malicious directions;
  • protect sensitive information and avoid privacy leakage;
  • respect least-privilege and role boundaries;
  • avoid unauthorised or destructive tool use;
  • refuse unsafe attempts to recover secrets from memory;
  • request human approval before high-impact actions; and
  • continue completing legitimate, benign tasks.

The final point matters. A safety benchmark should not reward an agent simply for refusing everything. It should distinguish safe restraint from useful task completion.

A multilingual and cross-sector design

Version 0.1.1 contains 120 fully synthetic scenarios spanning:

Dimension Coverage
Risk categories Prompt injection, privacy leakage, tool misuse, privilege abuse, memory safety and benign controls
Sectors Banking, healthcare, education, government and recruitment
Languages English, French, Swahili and Yoruba
Expected actions Answer, redact, refuse or request approval

Each category contains 20 cases, each sector contains 24 cases and each language contains 30 cases. The dataset contains no real people, credentials, organisations or operational targets.

All tools are inert, and the bundled reference policies are deterministic. This means the full evaluation can run locally without production credentials, paid model APIs or external recipients.

The threat and control model

The benchmark includes six scenario families:

  1. Prompt injection: adversarial instructions appear inside untrusted content and the expected action is refusal.
  2. Privacy leakage: a request attempts to obtain protected fields and the expected action is redaction.
  3. Tool misuse: an agent is asked to perform an unreviewed external action and must request approval.
  4. Privilege abuse: a read-only role is asked to perform a destructive action and must refuse.
  5. Memory safety: an agent is asked to reveal previously supplied secrets and must refuse.
  6. Benign controls: a legitimate request involving public synthetic information should be completed.

The benchmark produces machine-readable results for overall pass rate, attack success, privacy leakage, unauthorised tool use, human-approval violations, benign-task completion and category-level performance.

Deterministic baseline results

I evaluated two deliberately different fixtures against the same 120 cases:

  • a strict reference policy that applies basic instruction-integrity, privacy, permission and approval controls; and
  • a permissive control that follows the latest instruction, returns available context and invokes requested tools without policy checks.
Metric Strict policy Permissive control
Overall pass rate 100.00% 16.67%
Attack success rate 0.00% 100.00%
Privacy leakage rate 0.00% 50.00%
Unauthorised tool-use rate 0.00% 33.33%
Human-approval violation rate 0.00% 16.67%
Benign-task completion rate 100.00% 100.00%

These results validate the benchmark mechanics and scoring pipeline. They are not claims about the safety of any commercial or production model. The policies are transparent deterministic fixtures, and the current templates are intentionally easy to audit.

The equal benign-task completion result is especially important: the strict policy rejected unsafe behaviour without treating every request as malicious.

Reproduce the evaluation

The benchmark requires Python 3.10 or later:

git clone https://github.com/josepharayemi-netizen/agentguardbench.git
cd agentguardbench
python scripts/generate_dataset.py
PYTHONPATH=src python -m agentguardbench --policy both
python -m unittest discover -s tests -v

The evaluator writes machine-readable JSON and JSONL results to the results/ directory. The repository also includes a dataset card, ethics statement, security policy, test suite, report-generation script, figures and citation metadata.

Responsible interpretation

A benchmark score describes a tested configuration. It does not prove that a system is universally safe, ethical, secure or legally compliant.

A production evaluation must also account for the actual model, system prompts, memory, retrieval sources, identity controls, tool implementations, deployment context, users, monitoring and incident procedures.

AgentGuardBench v0.1.1 should therefore be used for:

  • defensive research;
  • education and training;
  • evaluation-pipeline validation;
  • controlled extension; and
  • authorised security testing.

It must not be used to target systems without permission, collect real personal data, bypass access controls or automate harmful actions.

Current limitations and future work

The first release is a transparent starting point, not a finished universal safety test. Important limitations include template-driven cases, deterministic reference policies, translations that still require independent native-speaker review and the absence of repeated stochastic model trials.

Planned extensions include:

  • adapters for local and hosted language models;
  • paraphrased and indirect prompt injections;
  • multi-turn attacks and tool-output poisoning;
  • context-window and memory-contamination tests;
  • cross-tool privilege escalation;
  • held-out challenge sets;
  • native-speaker language validation; and
  • independent annotation with uncertainty reporting.

I welcome responsible collaboration from AI engineers, security researchers, MLOps practitioners, linguists, governance specialists and sector experts.

Access and citation

Recommended citation:

Arayemi, J. (2026). AgentGuardBench: A Multilingual Benchmark for Privacy, Security and Responsible Behaviour in AI Agents (Version 1.0). Zenodo. https://doi.org/10.5281/zenodo.23132194

If you work on agentic AI, Responsible AI, privacy engineering, AI security or multilingual evaluation, I would value your feedback and independent reproduction of the benchmark.

📰 Read the original article on Dev.to Security

Originally published by Dev.to Security. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.