الباحثون

Jingyi He

المنشورات 1

نسخة أولية وصول مفتوح

RewardExplainer: Learning Reward Model Explanations from Counterfactual Preference Feedback

Jingyi He, Nier Wu, Shuang Liu وآخرون · 2026

Reward models (RMs) are a key component of large language model post-training, providing reward signals for subsequent reinforcement learning. However, conventional discriminative RMs typically output only scalar scores, making it difficult to identify the response behaviors associated with their scoring decisions. Exi …

المؤلفون المشاركون