Authors

Yue Zhang

Publications 7

Preprint Open access

WOVEN: Weaving Visual World Modeling into Multimodal LLMs

Zheyu Fan, Yue Zhang, Mingkai Deng et al. · 2026

Multimodal large language models (MLLMs) struggle with spatial, embodied, physical, and temporal reasoning. We hypothesize that these failures reflect a shared deficit in visual transition reasoning, and test whether this capability can serve as a shared training primitive, one that different models can learn from diff …

Preprint Open access

Ego2Act: Evaluating Goal-Directed Manipulation in Egocentric Video Generation

Video generation models are increasingly being explored as world simulators for embodied planning and learning. To do so effectively, these models must not only generate visually appealing frames, but also predict how environments dynamically evolve when executing goal-directed actions. While evaluating these capabilit …

Preprint Open access

Semifactual Credit-Augmented Policy Optimization

Reinforcement learning with verifiable rewards (RLVR) has improved the reasoning capabilities of large language models (LLMs), yet their predictions remain sensitive to task-irrelevant prompt features. We investigate this sensitivity through semifactual prompt interventions that preserve the underlying problem and its …

Co-authors