الباحثون

Zhongni Hou

المنشورات 1

نسخة أولية وصول مفتوح

On-Policy Parameter Update Direction Underlies Generalization in LLM Post-Training

Shufan Shen, Zhongni Hou, Junshu Sun وآخرون · 2026

The strong generalization performance of on-policy post-training paradigms has motivated studies of their parameter update behaviors. However, these studies treat the observed behaviors only as byproducts in on-policy training, overlooking their potential to serve as optimization principles for improving the generaliza …

المؤلفون المشاركون