الباحثون

Difan Zou

المنشورات 6

نسخة أولية وصول مفتوح

Understanding the Weight Averaging Mechanism in LLM Training for Post-Training Quantization

Hanzhang Wang, Tianqi Shen, Zonglin Liu وآخرون · 2026

Large language models (LLMs) are typically pretrained in high precision but increasingly deployed with low-precision post-training quantization (PTQ). Recent studies have shown that using weight averaging during pretraining can improve PTQ performance compared with learning-rate decay, suggesting that it might provide …

نسخة أولية وصول مفتوح

From Static to Dynamic: On-Policy Distillation from Image to Video Diffusion Models

Bingqing Jiang, Li Luo, Zichao Yu وآخرون · 2026

On-policy distillation (OPD) specializes pretrained video diffusion models through teacher supervision along the student's own generation trajectory. Although large video models are natural teachers, developing specialized video experts can require costly video data and training, while querying them incurs substantiall …

نسخة أولية وصول مفتوح

Understanding On-Policy Distillation: A Mechanistic Interpretability Perspective via Sparse Crosscoders

Zichao Yu, Qianshuo Ye, Xu Wang وآخرون · 2026

On-policy distillation (OPD) is a widely adopted post-training technique for LLM reasoning. It is commonly believed to transfer knowledge from a stronger teacher, yet what OPD actually distills into the student's internal representations remains unclear. We study this question with sparse crosscoders, which learn one f …

نسخة أولية وصول مفتوح

Beyond Token Scale: Chunk-Level Sparse Autoencoders for Reliable Semantic Feature Discovery

Xu Wang, Yifan Yang, TingHao YU وآخرون · 2026

Sparse autoencoders (SAEs) expose features that help us understand and steer language models, but faithful reconstruction does not guarantee informative concepts. Token-level objectives reward lexical and formatting details alongside semantic content, all competing for a limited sparse budget. We introduce a family of …

نسخة أولية وصول مفتوح

CoRe: Co-Evolving Reward Models for Mitigating Latent Reward Hacking in Video Diffusion Models

Zhaolong Su, Yujin Han, Feng Wang وآخرون · 2026

Latent reward models (LRMs) enable efficient alignment of video diffusion models by scoring intermediate states directly in latent space. However, we find that optimizing against a fixed latent reward rapidly leads to latent reward hacking: the predicted reward stays high while perceptual and motion quality deteriorate …

المؤلفون المشاركون