Preprint Open access
On-Policy Parameter Update Direction Underlies Generalization in LLM Post-Training
The strong generalization performance of on-policy post-training paradigms has motivated studies of their parameter update behaviors. However, these studies treat the observed behaviors only as byproducts in on-policy training, overlooking their potential to serve as optimization principles for improving the generaliza …