Abstract

Vision-language-action (VLA) models often struggle to generalize to skill combinations absent from their fine-tuning demonstrations, even when every constituent skill has been demonstrated. We focus on a vision shortcut as one failure mode: during fine-tuning, visual observations can serve as a proxy for the instruction, so a policy may execute a demonstrated combination associated with similar observations rather than the instructed combination. This motivates training with counterfactual pairs formed by holding a demonstration observation fixed while changing the instruction to specify an undemonstrated combination. These pairs, however, lack corresponding demonstrated action targets. Crucially, the currently required skill has already been demonstrated, but actions from those executions cannot serve as direct targets because the same skill can require different actions across observations. We propose CRAFT, which transfers supervision from demonstrated executions of the required skill to counterfactual pairs using skill representations that can be reused across executions of the same skill. Across three VLA models and two simulation benchmarks, CRAFT improves success on undemonstrated combinations while maintaining high success on demonstrated ones; it also improves compositional generalization on a real robot. Project website: https://taegeunyang.github.io/craft/

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Yang, T., Na, Y., Cho, Y., & Yoon, S. E. (2026). Same Scene, Different Task: Skill Alignment for Compositional Generalization in VLAs. https://omanscience.com/en/articles/same-scene-different-task-skill-alignment-for-compositional-generalization-in-vlas

MLA 9

Yang, Taegeun, et al. "Same Scene, Different Task: Skill Alignment for Compositional Generalization in VLAs." https://omanscience.com/en/articles/same-scene-different-task-skill-alignment-for-compositional-generalization-in-vlas.

Chicago (author–date)

Yang, Taegeun, Youngju Na, Yoonki Cho, and Sung-eui Yoon. 2026. "Same Scene, Different Task: Skill Alignment for Compositional Generalization in VLAs." https://omanscience.com/en/articles/same-scene-different-task-skill-alignment-for-compositional-generalization-in-vlas.

Harvard

Yang, T., Na, Y., Cho, Y. and Yoon, S. E. (2026) 'Same Scene, Different Task: Skill Alignment for Compositional Generalization in VLAs', Available at: https://omanscience.com/en/articles/same-scene-different-task-skill-alignment-for-compositional-generalization-in-vlas.

Vancouver

Yang T, Na Y, Cho Y, Yoon SE. Same Scene, Different Task: Skill Alignment for Compositional Generalization in VLAs. https://omanscience.com/en/articles/same-scene-different-task-skill-alignment-for-compositional-generalization-in-vlas

IEEE

T. Yang, Y. Na, Y. Cho, and S. E. Yoon, "Same Scene, Different Task: Skill Alignment for Compositional Generalization in VLAs," https://omanscience.com/en/articles/same-scene-different-task-skill-alignment-for-compositional-generalization-in-vlas.