نسخة أولية وصول مفتوح
Token-World: World Modeling in Vision-Language Model Token Space for Robot Manipulation
A common approach to world-model simulation for vision-language-action (VLA) systems is to predict future RGB observations and then re-encode them into policy inputs, introducing an indirect interface between simulation and downstream policy execution. We instead investigate whether world dynamics can be modeled in a c …