الباحثون

Yonggan Fu

المنشورات 3

نسخة أولية وصول مفتوح

When Do We Need On-Policy Distillation? Distilling on Offline Student Rollouts Is Often Better

Siyan Zhao, Yonggan Fu, Jindong Jiang وآخرون · 2026

On-policy distillation (OPD) has become increasingly popular for transferring teacher capabilities to student models. In this work, we ask a critical research question: Is on-policy sampling always beneficial for distilling arbitrary teacher-student pairs? We show that a simple alternative, Semi-OPD, which distills fro …

نسخة أولية وصول مفتوح

Large Language Continuous Diffusion Models

Zhihan Yang, Wei Guo, Jean-Marie Lemercier وآخرون · 2026

Despite the success of discrete diffusion language models (dLMs) for fast parallel decoding, their non-smooth, high-dimensional space hinders trajectory steering for reasoning and inference acceleration. To overcome this, we present Sigma, the first large-scale (3B/8B) continuous dLM built on steerable, low-dimensional …

نسخة أولية وصول مفتوح

Mid-Harness: Scaling Actions Between Model and Harness for Terminal Agents

Minki Kang, Ryo Hachiuma, Shaokun Zhang وآخرون · 2026

Terminal agents act through stochastic model generations, yet the ability to generate a useful action does not ensure its reliable execution. A poor command (e.g., wrong package install) can change the environment in ways that hinder subsequent progress, even when the model could generate a better alternative. We inves …

المؤلفون المشاركون