الملخص
As Large Language Models (LLMs) evolve into autonomous agents that alter real-world states, ensuring operational safety across multi-step workflows has become a critical challenge. While recent work has moved beyond single-turn evaluation toward multi-turn paradigms, key limitations persist: step-level methods treat actions in isolation, missing how risks accumulate, while trajectory-level evaluations operate post-hoc, offering no opportunity for timely intervention. To address these limitations, we formalize Decoupled Proactive Safety Monitoring along three dimensions: whether to intervene, when to intervene, and what the risk is. We introduce PASTABench, a benchmark of 1,139 multi-turn trajectories spanning 5 risk categories and 13 subcategories. We further propose the Optimal Intervention Window (OIW), anchored by annotated Earliest-Signal and Trigger turns, to quantify intervention timeliness. Evaluation of 16 LLMs reveals that proactive intervention remains largely unsolved, with the best model achieving only 40.74% optimal-timing interventions. Fine-grained diagnosis further uncovers pervasive lexical overfitting: competitive safety scores of smaller models mask keyword hypersensitivity rather than genuine risk comprehension, as their proactive capability largely collapses once hazard vocabulary is neutralized.
الكلمات المفتاحية
الموضوع
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Sun, J., Zhou, Y., Zhu, H., Wen, P., Zhou, J., Han, S., & Guo, Y. (2026). PASTABench: Proactive Assessment of Sequential Trajectories for Agent Safety. https://omanscience.com/ar/articles/pastabench-proactive-assessment-of-sequential-trajectories-for-agent-safety
MLA 9
Sun, Jiapeng, et al. "PASTABench: Proactive Assessment of Sequential Trajectories for Agent Safety." https://omanscience.com/ar/articles/pastabench-proactive-assessment-of-sequential-trajectories-for-agent-safety.
شيكاغو (المؤلف–التاريخ)
Sun, Jiapeng, Yujin Zhou, Han Zhu, Pengcheng Wen, Jiayi Zhou, Sirui Han, and Yike Guo. 2026. "PASTABench: Proactive Assessment of Sequential Trajectories for Agent Safety." https://omanscience.com/ar/articles/pastabench-proactive-assessment-of-sequential-trajectories-for-agent-safety.
هارفارد
Sun, J., Zhou, Y., Zhu, H., Wen, P., Zhou, J., Han, S. and Guo, Y. (2026) 'PASTABench: Proactive Assessment of Sequential Trajectories for Agent Safety', Available at: https://omanscience.com/ar/articles/pastabench-proactive-assessment-of-sequential-trajectories-for-agent-safety.
فانكوفر
Sun J, Zhou Y, Zhu H, Wen P, Zhou J, Han S, et al. PASTABench: Proactive Assessment of Sequential Trajectories for Agent Safety. https://omanscience.com/ar/articles/pastabench-proactive-assessment-of-sequential-trajectories-for-agent-safety
IEEE
J. Sun, Y. Zhou, H. Zhu, P. Wen, J. Zhou, S. Han, and Y. Guo, "PASTABench: Proactive Assessment of Sequential Trajectories for Agent Safety," https://omanscience.com/ar/articles/pastabench-proactive-assessment-of-sequential-trajectories-for-agent-safety.