الباحثون

Xinyue Tan

المنشورات 2

نسخة أولية وصول مفتوح

Train Ahead, Distill Back: Bootstrapping On-Policy Self-Distillation for Large Language Models

Zheng Zhang, Xinyue Tan, Lufei Li وآخرون · 2026

On-policy self-distillation (OPSD) improves large language models by letting a self-teacher with privileged information provide dense token-level supervision on the model's own trajectories. Yet existing methods typically construct the self-teacher from the current, initial, or slowly averaged policy state, leaving the …

نسخة أولية وصول مفتوح

From Judgment Quality to Downstream Utility: Rethinking LLM-as-a-Judge for Open-Ended Tasks

Zheng Zhang, Lufei Li, Xinyue Tan وآخرون · 2026

LLM-as-a-Judge is increasingly used to evaluate policy responses on open-ended tasks that lack ground-truth answers. Existing work often directly converts the resulting judgments into reward signals for policy training, paying limited attention to intrinsic judgment quality and largely restricting the use of Judges to …

المؤلفون المشاركون