نسخة أولية وصول مفتوح
Residual Visual Credit Optimization: Conserved Evidence Routing for Multimodal Reinforcement Learning
Reinforcement learning with verifiable rewards scales multimodal reasoning, but an outcome reward says how much a trajectory is worth, not how that value should be spread over the decisions that produced it. We introduce Residual Visual Credit Optimization (RVCO), which treats token credit as a conserved routing proble …