Preprint Open access
ForeAct3D: Policy-Grounded Future World Modeling for VLA Policies
Robots need to anticipate how their actions will change the world, since manipulation success hinges on the resulting contacts and object motions. However, existing Vision-Language-Action (VLA) policies that predict future observations from shared features leave the forecast decoupled from the actions the policy will a …