Abstract

LLM personalization aims to generate responses aligned with individual users' preferences and needs. User-specific rubrics make these expectations explicit, providing direct supervision on what a satisfactory answer should cover. Existing rubric-guided approaches, however, exploit such guidance only at a coarse granularity, either by using rubrics to supervise the prediction of relevant aspects for subsequent generation or by reducing aspect coverage to a single response-level reward for reinforcement learning. This leaves a gap between specifying what a personalized answer should contain and teaching the model how to generate it. To bridge this gap, we propose GRASP, a rubric-aware on-policy self-distillation framework for LLM personalization that turns user-specific rubric aspects into fine-grained, token-level supervision. Specifically, GRASP pairs a rubric-free student with a rubric-informed teacher that additionally receives the target user-specific rubrics. By aligning their next-token distributions along on-policy trajectories generated by the student, GRASP transfers the teacher's rubric-conditioned guidance into the student, translating user-specific semantic requirements into dense token-level supervision. Since rubric-informed teachers can still produce inadequate supervision, we further introduce Rubric-based Teacher Validation (RTV), which retains only instances where the teacher sufficiently covers the target aspects, improving both supervision quality and training efficiency. Experiments on the LaMP-QA benchmark for personalized question answering demonstrate that GRASP achieves state-of-the-art performance across multiple backbones, supporting the effectiveness of rubric-guided token-level supervision for personalization. To ensure reproducibility, our code is available at https://github.com/SnowCharmQ/GRASP.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Qiu, Y., Zhao, X., Wang, C., Yan, C., Zu, R., Zhang, W., Jiang, X., Cai, J., & Zhang, Y. (2026). Rubric-Aware On-Policy Self-Distillation for LLM Personalization. https://omanscience.com/en/articles/rubric-aware-on-policy-self-distillation-for-llm-personalization

MLA 9

Qiu, Yilun, et al. "Rubric-Aware On-Policy Self-Distillation for LLM Personalization." https://omanscience.com/en/articles/rubric-aware-on-policy-self-distillation-for-llm-personalization.

Chicago (author–date)

Qiu, Yilun, Xiaoyan Zhao, Chengbing Wang, Cilin Yan, Rui Zu, Wanyang Zhang, Xiaolong Jiang, Jiayin Cai, and Yang Zhang. 2026. "Rubric-Aware On-Policy Self-Distillation for LLM Personalization." https://omanscience.com/en/articles/rubric-aware-on-policy-self-distillation-for-llm-personalization.

Harvard

Qiu, Y., Zhao, X., Wang, C., Yan, C., Zu, R., Zhang, W., Jiang, X., Cai, J. and Zhang, Y. (2026) 'Rubric-Aware On-Policy Self-Distillation for LLM Personalization', Available at: https://omanscience.com/en/articles/rubric-aware-on-policy-self-distillation-for-llm-personalization.

Vancouver

Qiu Y, Zhao X, Wang C, Yan C, Zu R, Zhang W, et al. Rubric-Aware On-Policy Self-Distillation for LLM Personalization. https://omanscience.com/en/articles/rubric-aware-on-policy-self-distillation-for-llm-personalization

IEEE

Y. Qiu, X. Zhao, C. Wang, C. Yan, R. Zu, W. Zhang, X. Jiang, J. Cai, and Y. Zhang, "Rubric-Aware On-Policy Self-Distillation for LLM Personalization," https://omanscience.com/en/articles/rubric-aware-on-policy-self-distillation-for-llm-personalization.