الباحثون

Houfeng Wang

المنشورات 2

نسخة أولية وصول مفتوح

Learn the Directions, Normalize the Gains: Post-Training Normalization for LoRA

Zailong Tian, Yanzhe Chen, Zhuoheng Han وآخرون · 2026

While Low-Rank Adaptation (LoRA) enables efficient task specialization, its learned updates can compromise capabilities beyond the target task. We identify \textbf{adaptation imbalance}: a few singular directions dominate the trained update, leaving its performance sensitive to how gains are allocated. We argue that \t …

نسخة أولية وصول مفتوح

On-Policy Visual Evidence Distillation

Shaohang Wei, Feifan Song, Guangyue Peng وآخرون · 2026

Visual agents solve problems by interleaving reasoning with image operations, and on-policy distillation (OPD) provides guidance from a strong teacher on student-generated interaction trajectories. However, image operations change the evidence available for subsequent reasoning, so local errors in evidence acquisition …

المؤلفون المشاركون