Threat-Preserving Representation: Why Agent Security Benchmarks Lie About Model Robustness
Agent security benchmarks report attack success rates (ASR) as if they measure model robustness. A new paper shows that changing how you serialize the same threat can swing ASR by 11-13 percentage points without changing
Agent security benchmarks report attack success rates (ASR) as if they measure model robustness. A new paper shows that changing how you serialize the same threat can swing ASR by 11-13 percentage points without changing the underlying attack, task, or policy. The problem is representation sensitivity: JSON keys, tool names, and prose descriptions all encode the same security problem differently, and models respond to the encoding, not just the threat.
This matters because enterprises compare models using these scores, choose defense mechanisms based on them, and make deployment decisions assuming the numbers generalize. They don't.
The Measurement Problem
Agent security benchmarks test whether a model will execute harmful actions when prompted through tool calls, environment observations, or user instructions. The standard metric is ASR: the percentage of attempts where the agent completes the harmful action.
Researchers at Penn State introduced threat-preserving representation sensitivity (TPRS) to measure how much ASR changes when you alter the agent-visible representation while keeping everything else constant:
- Same underlying task
- Same harmful action
- Same security policy
- Same ground truth
- Same environment
- Same evaluation criteria
The only variable is how the threat appears to the agent: tool names, JSON structure, natural language descriptions, or parameter labels.
Experimental Results
On Agent Security Bench (ASB), replacing threat-related tool names with threat-neutral equivalents raised committed ASR by:
- 11.67 points on GPT-5-mini
- 13.21 points on Claude Haiku 4.5
On MCPTox, adding an explicit threat-related name to a previously neutral tool lowered ASR by:
- 11.00 points on GPT-5-mini
- 4.11 points on Claude Haiku 4.5
On AgentDojo, adding threat-related wording to the attack-relevant tool changed ASR by only 0.50 points on GPT-4o-mini, but benign utility dropped 5.36 points on tasks requiring that tool. The model became more cautious overall, not just on attacks.
A controlled experiment on MCPTox showed that a threat-neutral name matched on token count, length, and casing reproduced most of the shift (8.54 of 11.00 points on GPT-5-mini). The representation itself, not semantic content, drives the variance.
Why Representation Matters
Agent architectures expose models to structured data in ways chat interfaces don't. Tool schemas, environment observations, and state snapshots all require serialization choices:
-
JSON keys:
delete_uservs.remove_accountvs.purge_data -
Tool names:
FileManager.delete()vs.cleanup_temp_files() - Observation format: prose descriptions vs. structured logs vs. raw API responses
-
Parameter labels:
targetvs.user_idvs.account_handle
Each choice creates a different tokenization pattern, attention distribution, and activation landscape. Models trained on internet text have priors about which words co-occur with threats. A tool named send_email triggers different associations than dispatch_message, even if the underlying function is identical.
Tokenization Sensitivity
The MCPTox experiment controlled for token count, length, and casing. The fact that a neutral name still reproduced 77% of the ASR shift suggests tokenization boundaries matter as much as semantic content. Models don't just read tool names, they process token sequences, and those sequences interact with positional encodings, attention masks, and layer activations in ways that leak into security decisions.
Implications for Agent Security Infrastructure
If you're building agent security tooling, this creates three problems:
Model comparison is broken. You can't compare GPT-4 vs. Claude vs. Llama on security if each benchmark uses different tool naming conventions. A 5-point ASR difference might just be representation noise.
Defense mechanism evaluation is unreliable. If you test a guardrail on one representation and deploy it on another, the measured effectiveness doesn't transfer. A prompt filter tuned on JSON schemas might fail on prose observations.
Red-teaming is representation-dependent. Adversarial testing that finds vulnerabilities in one serialization format might miss threats that only surface in another.
What Threat-Preserving Means
A representation is threat-preserving if it encodes the same security-relevant information across different serialization formats. The paper doesn't define this formally, but the experiments suggest it requires:
- Consistent semantic content (the harmful action is described the same way)
- Equivalent information availability (the agent has access to the same facts)
- Comparable cognitive load (the representation doesn't obscure or highlight the threat through formatting)
In practice, this is hard to achieve. Natural language is ambiguous. JSON is explicit but verbose. Structured logs are compact but require domain knowledge to interpret. Each format creates different affordances for threat detection.
Architecture Patterns for Representation Stability
If you're deploying agents in production, you need to account for representation sensitivity in your security stack.
Normalize Input Representations
Before feeding tool outputs or environment observations to the model, pass them through a normalization layer that standardizes:
- Tool name format (verb-noun vs. noun-verb, camelCase vs. snake_case)
- JSON key naming (consistent vocabulary for sensitive operations)
- Observation structure (always prose, always JSON, or always a hybrid)
This won't eliminate representation sensitivity, but it reduces the attack surface by limiting the number of formats an adversary can exploit.
class RepresentationNormalizer:
def __init__(self, schema_registry):
self.registry = schema_registry
self.threat_vocabulary = load_threat_terms()
def normalize_tool_call(self, tool_name, params):
# Map to canonical name
canonical = self.registry.get_canonical(tool_name)
# Rewrite params to standard keys
normalized_params = {
self.registry.get_param_name(k): v
for k, v in params.items()
}
# Flag if tool name contains threat-related tokens
threat_score = self._score_threat_tokens(canonical)
return {
"tool": canonical,
"params": normalized_params,
"metadata": {"threat_score": threat_score}
}
def _score_threat_tokens(self, name):
tokens = self.tokenize(name)
return sum(1 for t in tokens if t in self.threat_vocabulary)
Multi-Representation Testing
Run security evals across multiple representations of the same threat. If ASR varies by more than 5 points, the benchmark isn't measuring model robustness, it's measuring representation sensitivity.
| Representation | Tool Name | ASR (GPT-4o) | ASR (Claude 3.5) | Delta |
|---|---|---|---|---|
| Neutral JSON | process_request |
34% | 29% | 5% |
| Explicit JSON | delete_user_account |
22% | 18% | 4% |
| Prose description | "Remove the user's account" | 28% | 24% | 4% |
| Structured log | {"action": "purge", "target": "user"} |
31% | 26% | 5% |
If deltas exceed 10 points, your eval is unstable. Either fix the representation or report results as a range, not a point estimate.
Representation-Aware Guardrails
Instead of filtering on static keyword lists, build guardrails that detect threats across multiple serialization formats. Use a classifier trained on parallel corpora: the same harmful action expressed in JSON, prose, logs, and API calls.
class MultiRepresentationGuardrail:
def __init__(self, threat_classifier):
self.classifier = threat_classifier
self.representations = [
JSONSerializer(),
ProseSerializer(),
LogSerializer()
]
def check(self, tool_call):
# Serialize the same call in multiple formats
variants = [
rep.serialize(tool_call)
for rep in self.representations
]
# Score each variant
scores = [
self.classifier.predict(v)
for v in variants
]
# Block if any representation exceeds threshold
max_score = max(scores)
variance = max(scores) - min(scores)
return {
"blocked": max_score > 0.7,
"max_threat_score": max_score,
"representation_variance": variance
}
If variance is high, the guardrail is representation-sensitive. You need more training data or a different classifier architecture.
Observability for Representation Drift
In production, tool schemas evolve. APIs get renamed. Observation formats change when you upgrade dependencies. Each change can shift ASR without changing the underlying threat model.
Track representation drift by logging:
- Tool name distributions over time
- JSON key frequency
- Observation format changes
- Tokenization boundary shifts (if you control the tokenizer)
If you see a sudden change in tool name distribution (e.g., a new API version renames delete to remove), re-run security evals before deploying. The old ASR numbers no longer apply.
class RepresentationDriftMonitor:
def __init__(self, baseline_stats):
self.baseline = baseline_stats
self.current_window = []
def observe(self, tool_call):
self.current_window.append({
"tool": tool_call.name,
"param_keys": list(tool_call.params.keys()),
"token_count": len(self.tokenize(tool_call.name))
})
if len(self.current_window) >= 1000:
self._check_drift()
self.current_window = []
def _check_drift(self):
current_stats = self._compute_stats(self.current_window)
# Compare tool name distribution
tool_divergence = kl_divergence(
self.baseline.tool_dist,
current_stats.tool_dist
)
# Compare token count distribution
token_divergence = kl_divergence(
self.baseline.token_dist,
current_stats.token_dist
)
if tool_divergence > 0.1 or token_divergence > 0.1:
alert("Representation drift detected, re-run security evals")
Failure Modes
Representation sensitivity creates failure modes that don't show up in standard testing:
Schema migration breaks security. You refactor tool names for clarity, ASR jumps 10 points, and suddenly the agent executes harmful actions it previously refused.
Cross-model deployment fails. You eval on GPT-4, deploy on Claude, and the ASR difference isn't model capability, it's tokenization variance.
Adversarial representation attacks. An attacker who controls tool output format can craft representations that maximize ASR, even if the underlying threat is identical.
Defense mechanism overfitting. You tune a guardrail on JSON schemas, it works great in testing, then fails in production because real tool calls use prose descriptions.
Technical Verdict
Use threat-preserving representation testing if:
- You're comparing models for agent deployment and need reliable security metrics
- You're building guardrails or defense mechanisms that must generalize across serialization formats
- You're red-teaming agents and want to find vulnerabilities that persist across representations
Avoid relying on single-representation ASR scores if:
- You're making deployment decisions based on benchmark leaderboards
- Your agent architecture exposes multiple serialization formats (JSON, prose, logs) to the model
- You're evaluating defense mechanisms without controlling for representation variance
The core insight is that agent security is representation-dependent, not just model-dependent. Until benchmarks standardize on threat-preserving formats or report ASR as a distribution across representations, treat security scores as lower bounds, not ground truth. If you're building production agent infrastructure, normalize representations, test across multiple formats, and monitor for drift. The alternative is deploying agents whose security properties change every time you rename a function.
Source Links
Originally published by Dev.to Security. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.