Abstract
Mathematical problem solving often requires deterministic computational steps that agents delegate to tools and implicitly trust. Yet tools can fail silently, returning plausible but incorrect results. How well can agents detect and correct corrupted tool call outputs? We study this through a controlled corruption framework where a hidden interceptor replaces tool call results with plausible incorrect information on targeted problems. We evaluate agents across 31 problems under four verification designs including no verification (baseline), mandatory same-context reflection, optional fresh-context verification, and optional structural verification. Without verification, corruption causes dramatic accuracy loss, from 100% down to 72.4%. Mandatory reflection fully recovers this performance to 100%. Optional verification improves accuracy only when models actively invoke it. Our results show that checking frequency is strongly associated with robustness differences, while unequal invocation prevents a controlled comparison of verifier quality. A supporting recovery experiment shows that full problem restart succeeds in 100% of cases after explicit detection. These findings demonstrate that verifier availability and verification policy are separate components of mathematical-agent reliability. Mandatory policies enforce verification while optional policies depend on the model's own choice to invoke it.
Keywords
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Jegatheesan, K., & Lihinikaduarachchi, G. (2026). When Tools Lie: Reliability of Mathematical Agents Under Corrupted Tool Feedback. https://omanscience.com/en/articles/when-tools-lie-reliability-of-mathematical-agents-under-corrupted-tool-feedback
MLA 9
Jegatheesan, Kavienan, and Gayathri Lihinikaduarachchi. "When Tools Lie: Reliability of Mathematical Agents Under Corrupted Tool Feedback." https://omanscience.com/en/articles/when-tools-lie-reliability-of-mathematical-agents-under-corrupted-tool-feedback.
Chicago (author–date)
Jegatheesan, Kavienan, and Gayathri Lihinikaduarachchi. 2026. "When Tools Lie: Reliability of Mathematical Agents Under Corrupted Tool Feedback." https://omanscience.com/en/articles/when-tools-lie-reliability-of-mathematical-agents-under-corrupted-tool-feedback.
Harvard
Jegatheesan, K. and Lihinikaduarachchi, G. (2026) 'When Tools Lie: Reliability of Mathematical Agents Under Corrupted Tool Feedback', Available at: https://omanscience.com/en/articles/when-tools-lie-reliability-of-mathematical-agents-under-corrupted-tool-feedback.
Vancouver
Jegatheesan K, Lihinikaduarachchi G. When Tools Lie: Reliability of Mathematical Agents Under Corrupted Tool Feedback. https://omanscience.com/en/articles/when-tools-lie-reliability-of-mathematical-agents-under-corrupted-tool-feedback
IEEE
K. Jegatheesan, and G. Lihinikaduarachchi, "When Tools Lie: Reliability of Mathematical Agents Under Corrupted Tool Feedback," https://omanscience.com/en/articles/when-tools-lie-reliability-of-mathematical-agents-under-corrupted-tool-feedback.