الباحثون

Jian Kang

المنشورات 2

نسخة أولية وصول مفتوح

Overcoming Prior Barriers: Supervised Fine-Tuning under Long-Tail Distribution

Haohui Wang, Jiahao Xu, Wangzhi Zhan وآخرون · 2026

Supervised fine-tuning (SFT) adapts pretrained large language models (LLMs) to downstream tasks, but the required concepts can receive substantially different levels of pretrained support. Frequent concepts are more likely to be well learned, whereas rare concepts may remain weakly represented. We introduce a novel not …

نسخة أولية وصول مفتوح

1% of Tokens Can Be Enough: On Gradient Estimation in On-Policy Distillation

Huanxin Sheng, Zhiling Ye, Haonan Wang وآخرون · 2026

Sparse on-policy distillation (OPD) allocates teacher supervision to a small subset of tokens in student-generated trajectories. However, useful teacher guidance can yield a noisy update when its gradient is estimated from a sampled next token. We study this estimation problem at a fixed prefix in information geometry …

المؤلفون المشاركون