الباحثون

Ilgee Hong

المنشورات 1

نسخة أولية وصول مفتوح

Training LLM Judges from Language Feedback via Position-Selective Self-Distillation

Ilgee Hong, Changlong Yu, Zhenghao Xu وآخرون · 2026

We study training LLM judges from natural language feedback, especially for subjective tasks where the verdict depends strongly on which evaluation criteria the judge invokes and how it weighs them. The dominant approach, outcome-supervised RL (e.g., GRPO), credits every token in the rollout with a single scalar determ …

المؤلفون المشاركون