الباحثون

Kun Kuang

المنشورات 2

نسخة أولية وصول مفتوح

MetaOPD: Meta-Learned Token Weighting for On-Policy Distillation

Zipeng Wang, Xinpeng Dong, Yuefan Wang وآخرون · 2026

On-policy distillation (OPD) trains a student on its own generated responses using token-level teacher supervision. However, uniform weighting overlooks differences in token learning value, while existing weighting methods rely on predefined mappings from prediction signals to token weights. These mappings are not lear …

نسخة أولية وصول مفتوح

Towards Efficient Reasoning: Learning Causal Shortcuts for Diffusion Language Models

Dian Jin, Kairong Han, Baohong Li وآخرون · 2026

Diffusion Language Models (DLMs) have attracted significant attention for their strong reasoning ability. However, under a bidirectional attention mechanism, DLMs operate over an exponentially large exploration space compared to autoregressive models (ARMs), making it challenging to focus on reasoning-guiding tokens un …

المؤلفون المشاركون