Authors

Zhe Wang

Publications 3

Preprint Open access

Adaptive Reward Routing: Dynamic Multi-Reward Optimization for Joint Audio-Video Diffusion via Forward-Process RL

Multi-reward guided reinforcement learning (i.e., RL) offers a promising way to improve joint audio-video diffusion models along several complementary objectives, including modality-specific quality, cross-modal semantic alignment, and temporal synchronization. Its effectiveness, however, depends on two quantities that …

Preprint Open access

Spotter: Let the Embodied Model Lead, and the VLM Reflect for It

Long Li, Qichao Zhao, Yue Yang et al. · 2026

Current embodied models do not respond to their own failures, although what just went wrong could inform a small adjustment on the next attempt, the kind of reflection behind the gains of thinking in language models. We test whether they can repair a known error, which requires producing a correction and judging whethe …

Co-authors