Abstract
Security benchmarks for LLM-based agents often report the attack success rate (ASR) as a measure of model robustness and use these scores to compare different models and defense mechanisms, assuming that they describe the security of the agent. In this paper, we explore whether it also influences the benchmark's measurement. To measure the effect of the benchmark representation, we introduce threat-preserving representation sensitivity (TPRS), which measures how much the ASR changes when we change the agent-visible representation while holding the underlying task, harmful action, security policy, ground truth, environment, and the evaluation criteria fixed. On Agent Security Bench (ASB), replacing threat-related tool names with threat-neutral names raises the committed attack success rate by 11.67 percentage points on GPT-5-mini and by 13.21 points on Claude Haiku 4.5. On MCPTox, replacing the original neutral tool name with an explicit threat-related name lowers the ASR by 11.00 percentage points on GPT-5-mini and 4.11 points on Claude Haiku 4.5. On AgentDojo, adding threat-related wording to the attack-relevant tool changes ASR by only 0.50 percentage points on GPT-4o-mini, yet the benign utility falls by 5.36 points on tasks requiring that tool. We ran an experiment on MCPTox where we observed that a threat-neutral name matched on token count, length, and casing reproduces most of the shift produced by the threat-explicit name (8.54 of 11.00 points on GPT-5-mini). The results show that a security score measured under one representation may fail to generalize across threat-preserving representations of the same security problem. Robustness claims should therefore be supported by performance across a controlled set of threat-preserving representations rather than relying on a single representation-dependent score.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Karamchandani, N., Nagasubramaniam, P., Xie, X., Zhu, S., & Wu, D. (2026). Threat-Preserving Representation Sensitivity in Agent-Security Benchmarks. https://omanscience.com/en/articles/threat-preserving-representation-sensitivity-in-agent-security-benchmarks
MLA 9
Karamchandani, Neeraj, et al. "Threat-Preserving Representation Sensitivity in Agent-Security Benchmarks." https://omanscience.com/en/articles/threat-preserving-representation-sensitivity-in-agent-security-benchmarks.
Chicago (author–date)
Karamchandani, Neeraj, Piyush Nagasubramaniam, Xinhong Xie, Sencun Zhu, and Dinghao Wu. 2026. "Threat-Preserving Representation Sensitivity in Agent-Security Benchmarks." https://omanscience.com/en/articles/threat-preserving-representation-sensitivity-in-agent-security-benchmarks.
Harvard
Karamchandani, N., Nagasubramaniam, P., Xie, X., Zhu, S. and Wu, D. (2026) 'Threat-Preserving Representation Sensitivity in Agent-Security Benchmarks', Available at: https://omanscience.com/en/articles/threat-preserving-representation-sensitivity-in-agent-security-benchmarks.
Vancouver
Karamchandani N, Nagasubramaniam P, Xie X, Zhu S, Wu D. Threat-Preserving Representation Sensitivity in Agent-Security Benchmarks. https://omanscience.com/en/articles/threat-preserving-representation-sensitivity-in-agent-security-benchmarks
IEEE
N. Karamchandani, P. Nagasubramaniam, X. Xie, S. Zhu, and D. Wu, "Threat-Preserving Representation Sensitivity in Agent-Security Benchmarks," https://omanscience.com/en/articles/threat-preserving-representation-sensitivity-in-agent-security-benchmarks.