Abstract
Latent reward models (LRMs) enable efficient alignment of video diffusion models by scoring intermediate states directly in latent space. However, we find that optimizing against a fixed latent reward rapidly leads to latent reward hacking: the predicted reward stays high while perceptual and motion quality deteriorate. Our analysis identifies distributional escape as the central cause: within a few hundred updates, the generator moves beyond the reward model's training support, where its scores no longer reflect video quality. Based on this insight, we introduce CoRe, a co-evolving reward framework that treats latent-space alignment as a dynamic interaction between the generator and the reward model. Rather than optimizing against a stationary proxy, CoRe continually refits the reward model on the generator's current samples while anchoring it to real-video preferences, so the generator cannot gain reward by drifting away from the data. On Wan2.1-T2V-1.3B, experiments show that CoRe consistently improves generation quality over both the pretrained model and prior alignment methods, while avoiding the quality collapse of fixed-reward optimization.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Su, Z., Han, Y., Wang, F., Dong, J., Hu, H., & Zou, D. (2026). CoRe: Co-Evolving Reward Models for Mitigating Latent Reward Hacking in Video Diffusion Models. https://omanscience.com/en/articles/core-co-evolving-reward-models-for-mitigating-latent-reward-hacking-in-video-diffusion-models
MLA 9
Su, Zhaolong, et al. "CoRe: Co-Evolving Reward Models for Mitigating Latent Reward Hacking in Video Diffusion Models." https://omanscience.com/en/articles/core-co-evolving-reward-models-for-mitigating-latent-reward-hacking-in-video-diffusion-models.
Chicago (author–date)
Su, Zhaolong, Yujin Han, Feng Wang, Jameson Dong, Hins Hu, and Difan Zou. 2026. "CoRe: Co-Evolving Reward Models for Mitigating Latent Reward Hacking in Video Diffusion Models." https://omanscience.com/en/articles/core-co-evolving-reward-models-for-mitigating-latent-reward-hacking-in-video-diffusion-models.
Harvard
Su, Z., Han, Y., Wang, F., Dong, J., Hu, H. and Zou, D. (2026) 'CoRe: Co-Evolving Reward Models for Mitigating Latent Reward Hacking in Video Diffusion Models', Available at: https://omanscience.com/en/articles/core-co-evolving-reward-models-for-mitigating-latent-reward-hacking-in-video-diffusion-models.
Vancouver
Su Z, Han Y, Wang F, Dong J, Hu H, Zou D. CoRe: Co-Evolving Reward Models for Mitigating Latent Reward Hacking in Video Diffusion Models. https://omanscience.com/en/articles/core-co-evolving-reward-models-for-mitigating-latent-reward-hacking-in-video-diffusion-models
IEEE
Z. Su, Y. Han, F. Wang, J. Dong, H. Hu, and D. Zou, "CoRe: Co-Evolving Reward Models for Mitigating Latent Reward Hacking in Video Diffusion Models," https://omanscience.com/en/articles/core-co-evolving-reward-models-for-mitigating-latent-reward-hacking-in-video-diffusion-models.