الباحثون

Shigang Chen

المنشورات 1

نسخة أولية وصول مفتوح

PR-OPD: Privileged Representation On-policy Self-Distillation for Agentic Reinforcement Learning

Muyang Li, Jie Yang, Zhengyu Fang وآخرون · 2026

Language-model agents are usually trained by reinforcement learning from one reward per episode, and privileged self-distillation enriches it by letting the same policy, given a skill, teach its skill-free self through token probabilities. However, we identify two phenomena that question this channel. Invisible Advanta …

المؤلفون المشاركون