الملخص

تمت ترجمة أجزاء من هذه الصفحة آلياً وقد تحتوي على أخطاء.

Tool-using language-model agents select and execute third-party artifacts. Different implementations can return the requested output while producing hidden execution effects that task-, attack-, or choice-based evaluations may miss. We study functional counterfeits: implementations that match benign alternatives on the requested output but add an effect forbidden by the task contract. We introduce SINGED (Source Integrity and the Nonidentifiability Gap in Execution Decisions for LLM Agents), a controlled benchmark covering five primary and two held-out task families. It varies displayed rank, evidence depth, decision policy, model release, and agent configuration, while task and process oracles verify the artifact and execution path. Across 7,549 audited trials, the randomized-rank study finds counterfeit execution in 45% (27/60) of rank-one trials and none at later ranks. Cross-candidate comparison eliminates shallow failures and reduces layered failures from 15.7% to 4.2%, but leaves dependency failures; its benefit is uncertain on unseen effects and public-package structures. Moreover, seven releases with no counterfeit executions when benign alternatives are available execute the counterfeit in 55/175 single-source cells after alternatives are removed. SINGED thus exposes a rank-, evidence-, and choice-sensitive outcome-to-execution gap: evaluation must connect correct outputs to execution paths.

الكلمات المفتاحية

الموضوع

بيانات النشر

المجلة
غير متاح
وصول مفتوح
وصول مفتوح أخضر

اقتبس هذه المقالة

APA 7

Xu, X., Liang, Z., Du, M., Xie, Q., Ye, Q., Li, Y., & Hu, H. (2026). SINGED: المخرجات الصحيحة لا تضمن التنفيذ الآمن في وكلاء LLM. https://omanscience.com/ar/articles/singed-correct-outputs-do-not-certify-safe-execution-in-llm-agents

MLA 9

Xu, Xiaoyu, et al. "SINGED: المخرجات الصحيحة لا تضمن التنفيذ الآمن في وكلاء LLM." https://omanscience.com/ar/articles/singed-correct-outputs-do-not-certify-safe-execution-in-llm-agents.

شيكاغو (المؤلف–التاريخ)

Xu, Xiaoyu, Zi Liang, Minxin Du, Qipeng Xie, Qingqing Ye, Yuyuan Li, and Haibo Hu. 2026. "SINGED: المخرجات الصحيحة لا تضمن التنفيذ الآمن في وكلاء LLM." https://omanscience.com/ar/articles/singed-correct-outputs-do-not-certify-safe-execution-in-llm-agents.

هارفارد

Xu, X., Liang, Z., Du, M., Xie, Q., Ye, Q., Li, Y. and Hu, H. (2026) 'SINGED: المخرجات الصحيحة لا تضمن التنفيذ الآمن في وكلاء LLM', Available at: https://omanscience.com/ar/articles/singed-correct-outputs-do-not-certify-safe-execution-in-llm-agents.

فانكوفر

Xu X, Liang Z, Du M, Xie Q, Ye Q, Li Y, et al. SINGED: المخرجات الصحيحة لا تضمن التنفيذ الآمن في وكلاء LLM. https://omanscience.com/ar/articles/singed-correct-outputs-do-not-certify-safe-execution-in-llm-agents

IEEE

X. Xu, Z. Liang, M. Du, Q. Xie, Q. Ye, Y. Li, and H. Hu, "SINGED: المخرجات الصحيحة لا تضمن التنفيذ الآمن في وكلاء LLM," https://omanscience.com/ar/articles/singed-correct-outputs-do-not-certify-safe-execution-in-llm-agents.