← Latest reporting

One agent-security score can change when only the threat’s name changes

A new preprint holds tasks and policies fixed while changing agent-visible wording. Procurement tests should add controlled representation variants before ranking models or defences.

AI Capability FrontierPolicy, Standards and Governance
A flat collage shows the same threat-shaped core inside several differently named wrappers, with uneven response strips and a separate benign-utility strip rather than a single ranking.
Conceptual illustration generated with AI under editorial direction; it does not depict a real event.

What happened

Researchers introduced threat-preserving representation sensitivity and reported double-digit attack-success shifts on two benchmark setups after changing tool names while holding the underlying security problem fixed.

Why it matters

If a score depends on naming, a single benchmark representation can overstate how well a model or defence generalises to local tools, schemas and interfaces.

A preprint submitted October 2 defines threat-preserving representation sensitivity: change the agent-visible representation while holding the task, harmful action, policy, ground truth, environment and evaluation fixed. On Agent Security Bench, neutralising threat-related tool names raised committed attack success by 11.67 percentage points for GPT-5-mini and 13.21 points for Claude Haiku 4.5. On MCPTox, adding an explicit threat-related name lowered attack success by 11.00 and 4.11 points respectively. On AgentDojo the attack shift was only 0.50 point, but benign utility fell 5.36 points.

Those are configuration-specific preprint results, not universal model rankings. One matched neutral name reproduced 8.54 of the 11.00-point MCPTox shift for GPT-5-mini, strengthening the representation explanation without proving its size elsewhere.

A companion EvoRiskBench preprint describes 450 adversarial tasks across six scenarios and nine model-harness combinations, verified with runtime traces and environment states. Its highest reported attack success was 68.44%, but the authors say artifacts will be released only after safety and reproducibility checks, limiting independent replication today.

Add a representation matrix

For every local attack case, create controlled variants of tool name, parameter label, ordering and threat salience while preserving permitted and harmful outcomes. Report the distribution and benign utility, not the best single score. Predefine which variation reflects a realistic local interface.

The counterargument is that variants inflate evaluation cost. Use a small factorial sample first; large score movement justifies expansion. Record failed benign tasks as carefully as successful attacks because a defence that merely disables useful tools is not robust. The immediate decision is to block procurement rankings based on one representation until the candidate passes a local sensitivity panel.