نسخة أولية وصول مفتوح
$R^2$-WAM: Repair-and-Reject Post-Training for World Action Models
World Action Models (WAMs) emerge as a promising foundation for policy refinement by predicting the consequences of sampled actions. However, visually plausible predictions can mislead policy refinement if they fail to reflect the input actions. To address this mismatch, we introduce $R^2$-WAM, a two-stage repair-and-r …