Preprint Open access
IronMan: Information-Constrained Video-Action Learning for Robot Manipulation
Video Action Models (VAMs) couple visual dynamics modeling with action generation for robot manipulation. However, video representations are not naturally suited to action generation, as exposing the action policy to excessive visual detail can impair its generalization ability. Therefore, we introduce IronMan (Informa …