Authors

Shuangyong Song

Publications 1

Preprint Open access

Quality Determines Direction, Length Shapes Magnitude: Length Control for Open-Ended Reinforcement Learning

Zijun Weng, Zhongan Bi, Xuanang Gao et al. · 2026

Reinforcement learning (RL) changes not only what language models say, but also how much they say, often increasing response length at the cost of token efficiency. Controlling this length growth is particularly challenging in open-ended RL because (i) response length is entangled with quality, (ii) open-ended tasks la …

Co-authors