Abstract

AI agents now hold spend authority and settle payments without per-action human confirmation. The resulting loss is often not a security failure: a counterparty with the correct domain, the correct settlement address and a genuinely delivered service can charge more than it should, and no check keyed on identity will see it. We present three artefacts for measuring and reducing that loss. First, a taxonomy of agentic commerce fraud that separates five observation levels (agent reasoning, wire, settlement rail, counterparty, principal) from the request-level and history-level evidence available at each, and records which levels can observe which attacks. Second, Agentic Commerce Bench (ACB), a benchmark of twenty fraud classes generated from production aggregates, 1,647 catalogued service operations and 1,068 settlements, of which six involve a counterparty that is exactly who it claims to be. Third, gordonguard, an open-source detector stack and offline harness with which an operator can audit an agent configuration, replay hostile counterparties without an account, and run the same detectors inline. Calibrating to a stated false-positive budget on clean training traffic gives a 6.5% clean flag rate, replicated across three independent generations, and leaves eight of twenty classes no better than chance. On the four classes a reasoning layer can observe, a widely used agent security scanner run over its jailbreak-detection panel scores zero on all four, while correctly scoring 1.0 on a jailbreak supplied as a control. A measured median payment of $0.007 places a hard constraint on deployment: one human review costs 143 times the value of the payment it examines.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Srivastava, A., & Paul, D. (2026). Agentic Commerce Bench: Measuring Fraud Detection for Agents That Spend Money. https://omanscience.com/en/articles/agentic-commerce-bench-measuring-fraud-detection-for-agents-that-spend-money

MLA 9

Srivastava, Ankit, and Debjyoti Paul. "Agentic Commerce Bench: Measuring Fraud Detection for Agents That Spend Money." https://omanscience.com/en/articles/agentic-commerce-bench-measuring-fraud-detection-for-agents-that-spend-money.

Chicago (author–date)

Srivastava, Ankit, and Debjyoti Paul. 2026. "Agentic Commerce Bench: Measuring Fraud Detection for Agents That Spend Money." https://omanscience.com/en/articles/agentic-commerce-bench-measuring-fraud-detection-for-agents-that-spend-money.

Harvard

Srivastava, A. and Paul, D. (2026) 'Agentic Commerce Bench: Measuring Fraud Detection for Agents That Spend Money', Available at: https://omanscience.com/en/articles/agentic-commerce-bench-measuring-fraud-detection-for-agents-that-spend-money.

Vancouver

Srivastava A, Paul D. Agentic Commerce Bench: Measuring Fraud Detection for Agents That Spend Money. https://omanscience.com/en/articles/agentic-commerce-bench-measuring-fraud-detection-for-agents-that-spend-money

IEEE

A. Srivastava, and D. Paul, "Agentic Commerce Bench: Measuring Fraud Detection for Agents That Spend Money," https://omanscience.com/en/articles/agentic-commerce-bench-measuring-fraud-detection-for-agents-that-spend-money.