الباحثون

Jindong Jiang

المنشورات 2

نسخة أولية وصول مفتوح

When Do We Need On-Policy Distillation? Distilling on Offline Student Rollouts Is Often Better

Siyan Zhao, Yonggan Fu, Jindong Jiang وآخرون · 2026

On-policy distillation (OPD) has become increasingly popular for transferring teacher capabilities to student models. In this work, we ask a critical research question: Is on-policy sampling always beneficial for distilling arbitrary teacher-student pairs? We show that a simple alternative, Semi-OPD, which distills fro …

نسخة أولية وصول مفتوح

Mid-Harness: Scaling Actions Between Model and Harness for Terminal Agents

Minki Kang, Ryo Hachiuma, Shaokun Zhang وآخرون · 2026

Terminal agents act through stochastic model generations, yet the ability to generate a useful action does not ensure its reliable execution. A poor command (e.g., wrong package install) can change the environment in ways that hinder subsequent progress, even when the model could generate a better alternative. We inves …

المؤلفون المشاركون