الباحثون

Haojin Wang

المنشورات 1

نسخة أولية وصول مفتوح

Learning from Teacher Continuations at Student States

Haojin Wang, Dylan Zhang, Huaibo Chen وآخرون · 2026

We present OLIVE (OnLine InterVEntion). At each iteration, the evolving student policy generates a new prefix, the teacher continues it autoregressively, and the student is updated using cross-entropy computed on the teacher-generated tokens. Each design choice targets a corresponding limitation of existing distillatio …

المؤلفون المشاركون