الباحثون

Zhuofeng Li

المنشورات 1

نسخة أولية وصول مفتوح

Self-Retrospection Distillation: Turning Post-hoc Experiences into Prior Foresight

Reinforcement learning with verifiable rewards (RLVR) turns agent experience into learning signals primarily through scalar outcome rewards after interaction. For group-relative objectives, however, this signal vanishes when all rollouts receive the same reward, even though their trajectories may reveal useful informat …

المؤلفون المشاركون