Abstract
Reward hacking occurs when policy optimization exploits a brittle reward interface or an overly permissive proxy objective, improving the training score without improving the underlying response quality. This phenomenon is amplified in group-relative policy optimization: an unsupported reward can shift the group baseline and alter the updates of other rollouts, while post-hoc or purely relative weighting cannot represent group-wide uncertainty. We propose \emph{Self-Tuned Anchored Reliability Group-Relative Policy Optimization} (STAR-GRPO), a reliability-first advantage estimator based on paired assessments of the same rollout. STAR separates the quality signal from its learning influence: score disagreement determines rollout reliability, relative reliability enters a self-tuned robust location--scale fit before group normalization, and absolute group reliability attenuates the resulting bounded advantage. The analysis establishes coordinate and second-moment bounds, characterizes exact centering through the weighted location equation, and gives reliability-dependent attenuation guarantees for outlying rewards. We evaluate STAR-GRPO in two complementary reward-hacking regimes. In token-interface exploitation, STAR prevents runaway optimization of the deployed-interface score while improving the canonical quality signal. In rubric-proxy overoptimization for medical reasoning, STAR improves independent semantic evaluation, narrows the proxy--judge discrepancy, and reduces overclaim while optimizing the same task proxy. Together, these results show that reliability-first normalization offers a principled way to limit unsupported reward influence on both group baselines and policy updates, while retaining the task reward as the optimization target.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Tian, W., Li, Z., Xu, X., Zou, M., Peng, Y., & Zhuang, F. (2026). STAR-GRPO: Canonical Anchoring and Reliability-First Advantages against Representation-Dependent Reward Hacking. https://omanscience.com/en/articles/star-grpo-canonical-anchoring-and-reliability-first-advantages-against-representation-dependent-reward-hacking
MLA 9
Tian, Wan, et al. "STAR-GRPO: Canonical Anchoring and Reliability-First Advantages against Representation-Dependent Reward Hacking." https://omanscience.com/en/articles/star-grpo-canonical-anchoring-and-reliability-first-advantages-against-representation-dependent-reward-hacking.
Chicago (author–date)
Tian, Wan, Zhongyi Li, Xiang Xu, Minhao Zou, Yijie Peng, and Fuzhen Zhuang. 2026. "STAR-GRPO: Canonical Anchoring and Reliability-First Advantages against Representation-Dependent Reward Hacking." https://omanscience.com/en/articles/star-grpo-canonical-anchoring-and-reliability-first-advantages-against-representation-dependent-reward-hacking.
Harvard
Tian, W., Li, Z., Xu, X., Zou, M., Peng, Y. and Zhuang, F. (2026) 'STAR-GRPO: Canonical Anchoring and Reliability-First Advantages against Representation-Dependent Reward Hacking', Available at: https://omanscience.com/en/articles/star-grpo-canonical-anchoring-and-reliability-first-advantages-against-representation-dependent-reward-hacking.
Vancouver
Tian W, Li Z, Xu X, Zou M, Peng Y, Zhuang F. STAR-GRPO: Canonical Anchoring and Reliability-First Advantages against Representation-Dependent Reward Hacking. https://omanscience.com/en/articles/star-grpo-canonical-anchoring-and-reliability-first-advantages-against-representation-dependent-reward-hacking
IEEE
W. Tian, Z. Li, X. Xu, M. Zou, Y. Peng, and F. Zhuang, "STAR-GRPO: Canonical Anchoring and Reliability-First Advantages against Representation-Dependent Reward Hacking," https://omanscience.com/en/articles/star-grpo-canonical-anchoring-and-reliability-first-advantages-against-representation-dependent-reward-hacking.