Dev.to Security 🔐 Cybersecurity 👁 0 📖 7 min read

Threat-Preserving Representation: Why Agent Security Benchmarks Lie About Model Robustness

Agent security benchmarks report attack success rates (ASR) as if they measure model robustness. A new paper shows that changing how you serialize the same threat can swing ASR by 11-13 percentage points without changing

Agent security benchmarks report attack success rates (ASR) as if they measure model robustness. A new paper shows that changing how you serialize the same threat can swing ASR by 11-13 percentage points without changing the underlying attack, task, or policy. The problem is representation sensitivity: JSON keys, tool names, and prose descriptions all encode the same security problem differently, and models respond to the encoding, not just the threat.

This matters because enterprises compare models using these scores, choose defense mechanisms based on them, and make deployment decisions assuming the numbers generalize. They don't.

The Measurement Problem

Agent security benchmarks test whether a model will execute harmful actions when prompted through tool calls, environment observations, or user instructions. The standard metric is ASR: the percentage of attempts where the agent completes the harmful action.

Researchers at Penn State introduced threat-preserving representation sensitivity (TPRS) to measure how much ASR changes when you alter the agent-visible representation while keeping everything else constant:

  • Same underlying task
  • Same harmful action
  • Same security policy
  • Same ground truth
  • Same environment
  • Same evaluation criteria

The only variable is how the threat appears to the agent: tool names, JSON structure, natural language descriptions, or parameter labels.

Experimental Results

On Agent Security Bench (ASB), replacing threat-related tool names with threat-neutral equivalents raised committed ASR by:

  • 11.67 points on GPT-5-mini
  • 13.21 points on Claude Haiku 4.5

On MCPTox, adding an explicit threat-related name to a previously neutral tool lowered ASR by:

  • 11.00 points on GPT-5-mini
  • 4.11 points on Claude Haiku 4.5

On AgentDojo, adding threat-related wording to the attack-relevant tool changed ASR by only 0.50 points on GPT-4o-mini, but benign utility dropped 5.36 points on tasks requiring that tool. The model became more cautious overall, not just on attacks.

A controlled experiment on MCPTox showed that a threat-neutral name matched on token count, length, and casing reproduced most of the shift (8.54 of 11.00 points on GPT-5-mini). The representation itself, not semantic content, drives the variance.

Why Representation Matters

Agent architectures expose models to structured data in ways chat interfaces don't. Tool schemas, environment observations, and state snapshots all require serialization choices:

  • JSON keys: delete_user vs. remove_account vs. purge_data
  • Tool names: FileManager.delete() vs. cleanup_temp_files()
  • Observation format: prose descriptions vs. structured logs vs. raw API responses
  • Parameter labels: target vs. user_id vs. account_handle

Each choice creates a different tokenization pattern, attention distribution, and activation landscape. Models trained on internet text have priors about which words co-occur with threats. A tool named send_email triggers different associations than dispatch_message, even if the underlying function is identical.

Tokenization Sensitivity

The MCPTox experiment controlled for token count, length, and casing. The fact that a neutral name still reproduced 77% of the ASR shift suggests tokenization boundaries matter as much as semantic content. Models don't just read tool names, they process token sequences, and those sequences interact with positional encodings, attention masks, and layer activations in ways that leak into security decisions.

Implications for Agent Security Infrastructure

If you're building agent security tooling, this creates three problems:

  1. Model comparison is broken. You can't compare GPT-4 vs. Claude vs. Llama on security if each benchmark uses different tool naming conventions. A 5-point ASR difference might just be representation noise.

  2. Defense mechanism evaluation is unreliable. If you test a guardrail on one representation and deploy it on another, the measured effectiveness doesn't transfer. A prompt filter tuned on JSON schemas might fail on prose observations.

  3. Red-teaming is representation-dependent. Adversarial testing that finds vulnerabilities in one serialization format might miss threats that only surface in another.

What Threat-Preserving Means

A representation is threat-preserving if it encodes the same security-relevant information across different serialization formats. The paper doesn't define this formally, but the experiments suggest it requires:

  • Consistent semantic content (the harmful action is described the same way)
  • Equivalent information availability (the agent has access to the same facts)
  • Comparable cognitive load (the representation doesn't obscure or highlight the threat through formatting)

In practice, this is hard to achieve. Natural language is ambiguous. JSON is explicit but verbose. Structured logs are compact but require domain knowledge to interpret. Each format creates different affordances for threat detection.

Architecture Patterns for Representation Stability

If you're deploying agents in production, you need to account for representation sensitivity in your security stack.

Normalize Input Representations

Before feeding tool outputs or environment observations to the model, pass them through a normalization layer that standardizes:

  • Tool name format (verb-noun vs. noun-verb, camelCase vs. snake_case)
  • JSON key naming (consistent vocabulary for sensitive operations)
  • Observation structure (always prose, always JSON, or always a hybrid)

This won't eliminate representation sensitivity, but it reduces the attack surface by limiting the number of formats an adversary can exploit.

class RepresentationNormalizer:
    def __init__(self, schema_registry):
        self.registry = schema_registry
        self.threat_vocabulary = load_threat_terms()

    def normalize_tool_call(self, tool_name, params):
        # Map to canonical name
        canonical = self.registry.get_canonical(tool_name)

        # Rewrite params to standard keys
        normalized_params = {
            self.registry.get_param_name(k): v
            for k, v in params.items()
        }

        # Flag if tool name contains threat-related tokens
        threat_score = self._score_threat_tokens(canonical)

        return {
            "tool": canonical,
            "params": normalized_params,
            "metadata": {"threat_score": threat_score}
        }

    def _score_threat_tokens(self, name):
        tokens = self.tokenize(name)
        return sum(1 for t in tokens if t in self.threat_vocabulary)

Multi-Representation Testing

Run security evals across multiple representations of the same threat. If ASR varies by more than 5 points, the benchmark isn't measuring model robustness, it's measuring representation sensitivity.

Representation Tool Name ASR (GPT-4o) ASR (Claude 3.5) Delta
Neutral JSON process_request 34% 29% 5%
Explicit JSON delete_user_account 22% 18% 4%
Prose description "Remove the user's account" 28% 24% 4%
Structured log {"action": "purge", "target": "user"} 31% 26% 5%

If deltas exceed 10 points, your eval is unstable. Either fix the representation or report results as a range, not a point estimate.

Representation-Aware Guardrails

Instead of filtering on static keyword lists, build guardrails that detect threats across multiple serialization formats. Use a classifier trained on parallel corpora: the same harmful action expressed in JSON, prose, logs, and API calls.

class MultiRepresentationGuardrail:
    def __init__(self, threat_classifier):
        self.classifier = threat_classifier
        self.representations = [
            JSONSerializer(),
            ProseSerializer(),
            LogSerializer()
        ]

    def check(self, tool_call):
        # Serialize the same call in multiple formats
        variants = [
            rep.serialize(tool_call)
            for rep in self.representations
        ]

        # Score each variant
        scores = [
            self.classifier.predict(v)
            for v in variants
        ]

        # Block if any representation exceeds threshold
        max_score = max(scores)
        variance = max(scores) - min(scores)

        return {
            "blocked": max_score > 0.7,
            "max_threat_score": max_score,
            "representation_variance": variance
        }

If variance is high, the guardrail is representation-sensitive. You need more training data or a different classifier architecture.

Observability for Representation Drift

In production, tool schemas evolve. APIs get renamed. Observation formats change when you upgrade dependencies. Each change can shift ASR without changing the underlying threat model.

Track representation drift by logging:

  • Tool name distributions over time
  • JSON key frequency
  • Observation format changes
  • Tokenization boundary shifts (if you control the tokenizer)

If you see a sudden change in tool name distribution (e.g., a new API version renames delete to remove), re-run security evals before deploying. The old ASR numbers no longer apply.

class RepresentationDriftMonitor:
    def __init__(self, baseline_stats):
        self.baseline = baseline_stats
        self.current_window = []

    def observe(self, tool_call):
        self.current_window.append({
            "tool": tool_call.name,
            "param_keys": list(tool_call.params.keys()),
            "token_count": len(self.tokenize(tool_call.name))
        })

        if len(self.current_window) >= 1000:
            self._check_drift()
            self.current_window = []

    def _check_drift(self):
        current_stats = self._compute_stats(self.current_window)

        # Compare tool name distribution
        tool_divergence = kl_divergence(
            self.baseline.tool_dist,
            current_stats.tool_dist
        )

        # Compare token count distribution
        token_divergence = kl_divergence(
            self.baseline.token_dist,
            current_stats.token_dist
        )

        if tool_divergence > 0.1 or token_divergence > 0.1:
            alert("Representation drift detected, re-run security evals")

Failure Modes

Representation sensitivity creates failure modes that don't show up in standard testing:

  1. Schema migration breaks security. You refactor tool names for clarity, ASR jumps 10 points, and suddenly the agent executes harmful actions it previously refused.

  2. Cross-model deployment fails. You eval on GPT-4, deploy on Claude, and the ASR difference isn't model capability, it's tokenization variance.

  3. Adversarial representation attacks. An attacker who controls tool output format can craft representations that maximize ASR, even if the underlying threat is identical.

  4. Defense mechanism overfitting. You tune a guardrail on JSON schemas, it works great in testing, then fails in production because real tool calls use prose descriptions.

Technical Verdict

Use threat-preserving representation testing if:

  • You're comparing models for agent deployment and need reliable security metrics
  • You're building guardrails or defense mechanisms that must generalize across serialization formats
  • You're red-teaming agents and want to find vulnerabilities that persist across representations

Avoid relying on single-representation ASR scores if:

  • You're making deployment decisions based on benchmark leaderboards
  • Your agent architecture exposes multiple serialization formats (JSON, prose, logs) to the model
  • You're evaluating defense mechanisms without controlling for representation variance

The core insight is that agent security is representation-dependent, not just model-dependent. Until benchmarks standardize on threat-preserving formats or report ASR as a distribution across representations, treat security scores as lower bounds, not ground truth. If you're building production agent infrastructure, normalize representations, test across multiple formats, and monitor for drift. The alternative is deploying agents whose security properties change every time you rename a function.

Source Links

📰 Read the original article on Dev.to Security

Originally published by Dev.to Security. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.