الملخص
Modern embodied agents achieve impressive success rates, yet their actual instruction-following ability is far weaker than these numbers suggest. We trace this illusion to a structural property we term low scene entropy: when a visual scene admits only one valid task, language becomes redundant and a policy can score highly while barely using it. We introduce RoboFollow, a diagnostic benchmark with three principles: (1) High Scene Entropy: each training scene supports multiple kinematically distinct task branches, making vision alone insufficient and forcing reliance on language. (2) Hierarchical Diagnostic Protocol: a four-level protocol (L0--L3) progressively perturbs visual layout and semantics, probing whether equivalent instructions yield consistent behavior and distinct ones yield discriminable behavior across spatial relations, attributes, trajectory constraints, and logic. (3) Confound-Controlled Diagnosis: we simplify interaction objects, restrict actions to the trained repertoire and report stage-wise Intent and Execution scores, isolating comprehension from motor execution. Evaluation of nine VLA and WAM policies shows that strong L0 performance, where attained, does not reliably transfer to L1--L3 under our fine-tuning setup. Representative mitigations, including stronger VLM backbones, QA co-training, LangForce, and Classifier-Free Guidance, all fail to close this gap. RoboFollow exposes genuine instruction following as a critical, overlooked bottleneck. Code and dataset are available at https://github.com/AutoLab-SAI-SJTU/RoboFollow and https://huggingface.co/datasets/AutoLab-SJTU/robofollow-data.
الكلمات المفتاحية
الموضوع
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Guo, C., Xie, Y., Tan, B., Chang, Z., Yin, Z., Ma, Q., Wang, Y., Liang, C., & Zhang, Z. (2026). RoboFollow: Unveiling the Instruction Following Mirage in Embodied Agents. https://omanscience.com/ar/articles/robofollow-unveiling-the-instruction-following-mirage-in-embodied-agents
MLA 9
Guo, Chang, et al. "RoboFollow: Unveiling the Instruction Following Mirage in Embodied Agents." https://omanscience.com/ar/articles/robofollow-unveiling-the-instruction-following-mirage-in-embodied-agents.
شيكاغو (المؤلف–التاريخ)
Guo, Chang, Yukun Xie, Bohan Tan, Zheng Chang, Zhaokai Yin, Qianli Ma, Yingqiao Wang, Chao Liang, and Zhipeng Zhang. 2026. "RoboFollow: Unveiling the Instruction Following Mirage in Embodied Agents." https://omanscience.com/ar/articles/robofollow-unveiling-the-instruction-following-mirage-in-embodied-agents.
هارفارد
Guo, C., Xie, Y., Tan, B., Chang, Z., Yin, Z., Ma, Q., Wang, Y., Liang, C. and Zhang, Z. (2026) 'RoboFollow: Unveiling the Instruction Following Mirage in Embodied Agents', Available at: https://omanscience.com/ar/articles/robofollow-unveiling-the-instruction-following-mirage-in-embodied-agents.
فانكوفر
Guo C, Xie Y, Tan B, Chang Z, Yin Z, Ma Q, et al. RoboFollow: Unveiling the Instruction Following Mirage in Embodied Agents. https://omanscience.com/ar/articles/robofollow-unveiling-the-instruction-following-mirage-in-embodied-agents
IEEE
C. Guo, Y. Xie, B. Tan, Z. Chang, Z. Yin, Q. Ma, Y. Wang, C. Liang, and Z. Zhang, "RoboFollow: Unveiling the Instruction Following Mirage in Embodied Agents," https://omanscience.com/ar/articles/robofollow-unveiling-the-instruction-following-mirage-in-embodied-agents.