الملخص

While Vision-Language-Action (VLA) models perform strongly on manipulation tasks, their responses to invalid task premises remain underexplored. Existing evaluations of premise conflicts often focus on terminal task outcomes, yet task failure alone cannot distinguish behavioral disengagement from continued pursuit followed by an execution error. We call the latter pattern Failed Persistence. To study this phenomenon, we introduce ConflictVLA-Bench, which pairs conflict rollouts with premise-consistent reference rollouts and evaluates both outcomes and execution processes. Built on LIBERO, the benchmark contains 2,826 prompt-conditioned conflict tasks spanning four conflict families, four structural configurations, and two prompt conditions. Across all eight VLAs, invalid premises reduce original goal completion by at least 17.3 percentage points, with the reduction reaching 56.2 percentage points for OpenVLA. Crucially, even when models succeed on premise-consistent tasks and fail on their matched conflict tasks, they often continue to approach the original targets, retain early trajectory structure, and show limited action magnitude suppression. Failed Persistence therefore recurs across the evaluated models. Explicit premise checking does not consistently produce selective and coordinated behavioral changes. These findings show that terminal failure alone establishes neither behavioral disengagement nor refusal and that outcomes alone are insufficient for VLA evaluation. Experimental data and additional details are available on the project page: https://github.com/EmbodiedAISurvey/ConflictVLA-Bench

الكلمات المفتاحية

الموضوع

بيانات النشر

المجلة
غير متاح
وصول مفتوح
وصول مفتوح أخضر

اقتبس هذه المقالة

APA 7

Hou, L., Wu, Y., & Chang, Y. (2026). ConflictVLA-Bench: Benchmarking Behavioral Responses of Vision-Language-Action Models to Premise Conflicts. https://omanscience.com/ar/articles/conflictvla-bench-benchmarking-behavioral-responses-of-vision-language-action-models-to-premise-conflicts

MLA 9

Hou, Liyu, et al. "ConflictVLA-Bench: Benchmarking Behavioral Responses of Vision-Language-Action Models to Premise Conflicts." https://omanscience.com/ar/articles/conflictvla-bench-benchmarking-behavioral-responses-of-vision-language-action-models-to-premise-conflicts.

شيكاغو (المؤلف–التاريخ)

Hou, Liyu, Yuan Wu, and Yi Chang. 2026. "ConflictVLA-Bench: Benchmarking Behavioral Responses of Vision-Language-Action Models to Premise Conflicts." https://omanscience.com/ar/articles/conflictvla-bench-benchmarking-behavioral-responses-of-vision-language-action-models-to-premise-conflicts.

هارفارد

Hou, L., Wu, Y. and Chang, Y. (2026) 'ConflictVLA-Bench: Benchmarking Behavioral Responses of Vision-Language-Action Models to Premise Conflicts', Available at: https://omanscience.com/ar/articles/conflictvla-bench-benchmarking-behavioral-responses-of-vision-language-action-models-to-premise-conflicts.

فانكوفر

Hou L, Wu Y, Chang Y. ConflictVLA-Bench: Benchmarking Behavioral Responses of Vision-Language-Action Models to Premise Conflicts. https://omanscience.com/ar/articles/conflictvla-bench-benchmarking-behavioral-responses-of-vision-language-action-models-to-premise-conflicts

IEEE

L. Hou, Y. Wu, and Y. Chang, "ConflictVLA-Bench: Benchmarking Behavioral Responses of Vision-Language-Action Models to Premise Conflicts," https://omanscience.com/ar/articles/conflictvla-bench-benchmarking-behavioral-responses-of-vision-language-action-models-to-premise-conflicts.