Abstract

The strong generalization performance of on-policy post-training paradigms has motivated studies of their parameter update behaviors. However, these studies treat the observed behaviors only as byproducts in on-policy training, overlooking their potential to serve as optimization principles for improving the generalization of other paradigms such as supervised fine-tuning (SFT). To address this limitation, we investigate whether there exists a specific on-policy update behavior that can achieve such improvements. First, our analyses reveal that SFT updates parameters along consistent directions, while the on-policy paradigm continuously adjusts the direction during training. This difference inspires us to focus on the cumulative update direction of each parameter as a promising behavior. Then, we evaluate its effectiveness for improving generalization by proposing On-Policy direction-constrained Supervised Fine-Tuning (OPSFT), which constrains SFT updates to the direction identified by on-policy paradigms. The strong performance of OPSFT indicates that the generalization advantage of on-policy paradigms can be transferred to SFT through the parameter update direction. Once such a direction is identified, even SFT can generalize with its updates constrained to this direction. This finding offers two practical benefits by combining the strong generalization of on-policy paradigms with the advantages of SFT, including the high training efficiency and ability to leverage high-quality trajectories. For efficiency, we identify update directions that support strong generalization using a few on-policy training steps, and subsequently apply OPSFT to achieve high training efficiency. For leveraging high-quality trajectories, OPSFT can utilize these trajectories to continue improving a post-trained model along its update direction without disrupting the ability learned from on-policy training.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Shen, S., Hou, Z., Sun, J., Zhang, Y., Lin, W., Yin, G., Huang, Q., & Wang, S. (2026). On-Policy Parameter Update Direction Underlies Generalization in LLM Post-Training. https://omanscience.com/en/articles/on-policy-parameter-update-direction-underlies-generalization-in-llm-post-training

MLA 9

Shen, Shufan, et al. "On-Policy Parameter Update Direction Underlies Generalization in LLM Post-Training." https://omanscience.com/en/articles/on-policy-parameter-update-direction-underlies-generalization-in-llm-post-training.

Chicago (author–date)

Shen, Shufan, Zhongni Hou, Junshu Sun, Yufei Zhang, Wei Lin, Guojun Yin, Qingming Huang, and Shuhui Wang. 2026. "On-Policy Parameter Update Direction Underlies Generalization in LLM Post-Training." https://omanscience.com/en/articles/on-policy-parameter-update-direction-underlies-generalization-in-llm-post-training.

Harvard

Shen, S., Hou, Z., Sun, J., Zhang, Y., Lin, W., Yin, G., Huang, Q. and Wang, S. (2026) 'On-Policy Parameter Update Direction Underlies Generalization in LLM Post-Training', Available at: https://omanscience.com/en/articles/on-policy-parameter-update-direction-underlies-generalization-in-llm-post-training.

Vancouver

Shen S, Hou Z, Sun J, Zhang Y, Lin W, Yin G, et al. On-Policy Parameter Update Direction Underlies Generalization in LLM Post-Training. https://omanscience.com/en/articles/on-policy-parameter-update-direction-underlies-generalization-in-llm-post-training

IEEE

S. Shen, Z. Hou, J. Sun, Y. Zhang, W. Lin, G. Yin, Q. Huang, and S. Wang, "On-Policy Parameter Update Direction Underlies Generalization in LLM Post-Training," https://omanscience.com/en/articles/on-policy-parameter-update-direction-underlies-generalization-in-llm-post-training.