نسخة أولية وصول مفتوح
Measuring and Mitigating Solution Mode Collapse in RLVR
A language model (LM) can usually answer the same question in more than one way, but reinforcement learning with verifiable rewards (RLVR) is indifferent to which correct answer a model produces. A solution will earn the same reward whether it is the thousandth copy of a familiar answer or one the model has never produ …