Authors

Xuanang Gao

Publications 2

Preprint Open access

Visual sensitivity is not claim retractability: persistence-aware credit assignment for multimodal reinforcement learning

Zhongan Bi, Kepeng Lin, Xuanang Gao et al. · 2026

Reinforcement Learning with Verifiable Rewards (RLVR) has been extended to Large Vision-Language Models (LVLMs), and perception-aware methods further encourage policies to rely on visual evidence. Yet relying on the image does not guarantee that visual claims are supported by it. Before RL training, 27.81% of the corre …

Preprint Open access

Quality Determines Direction, Length Shapes Magnitude: Length Control for Open-Ended Reinforcement Learning

Zijun Weng, Zhongan Bi, Xuanang Gao et al. · 2026

Reinforcement learning (RL) changes not only what language models say, but also how much they say, often increasing response length at the cost of token efficiency. Controlling this length growth is particularly challenging in open-ended RL because (i) response length is entangled with quality, (ii) open-ended tasks la …

Co-authors