الباحثون

Shuang Liu

المنشورات 2

نسخة أولية وصول مفتوح

Look Before You Select: Rethinking Vocabulary Sparsification in On-Policy Distillation

Yongliang Miao, Shuang Liu, Yanguang Liu وآخرون · 2026

On-policy distillation (OPD) uses teacher correction on student-generated responses. Full-vocabulary correction can provide important corrections even for tokens that the student assigns low probability, but backpropagating through all token logits becomes memory-intensive for long sequences. Existing memory-saving app …

نسخة أولية وصول مفتوح

RewardExplainer: Learning Reward Model Explanations from Counterfactual Preference Feedback

Jingyi He, Nier Wu, Shuang Liu وآخرون · 2026

Reward models (RMs) are a key component of large language model post-training, providing reward signals for subsequent reinforcement learning. However, conventional discriminative RMs typically output only scalar scores, making it difficult to identify the response behaviors associated with their scoring decisions. Exi …

المؤلفون المشاركون