Abstract

Generative reward models (GRMs) are important for LLM optimization. Unlike scalar reward models, GRMs generate natural-language critiques alongside preference judgments, providing finer-grained evaluation signals. Their effectiveness depends heavily on critique reliability. However, existing GRM training typically uses final preference correctness as outcome supervision. Because the preference outcome space is highly constrained, unreliable critiques can still yield correct outcomes and thus be reinforced. Recent work leverages human critiques for process supervision, but such critiques are scarce and are often reduced to scalar rewards, leaving their fine-grained evaluative information underutilized. We argue that evaluative criteria learned from human critiques can be generalized to broader outcome-only preference data. To this end, we propose \textbf{EnGRICH}, a GRM training framework that pairs the GRM with a training-time MetaCritic learned from a small set of human critiques. MetaCritic constructs response-specific rubrics and uses them to evaluate the evidence coverage and correctness of generated critiques. The resulting signals provide both process rewards for fine-grained credit assignment and structured guidance for exploring better critiques. During GRM training, MetaCritic is further optimized to generalize human-grounded evaluative criteria to outcome-only data. At inference, the trained GRM operates independently. Experiments across seven reward-model benchmarks show that EnGRICH consistently improves over competitive baselines, while further analyses validate the effectiveness of its core mechanisms.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Li, X., Wang, B., Li, H., Wang, H., Zhou, Y., Pan, Q., Chen, B., Liu, Y., Zhang, M., & Ai, Q. (2026). EnGRICH: Enhancing Generative Reward Modeling with Critiques from Humans. https://omanscience.com/en/articles/engrich-enhancing-generative-reward-modeling-with-critiques-from-humans

MLA 9

Li, Xuancheng, et al. "EnGRICH: Enhancing Generative Reward Modeling with Critiques from Humans." https://omanscience.com/en/articles/engrich-enhancing-generative-reward-modeling-with-critiques-from-humans.

Chicago (author–date)

Li, Xuancheng, Beining Wang, Haitao Li, Heng Wang, Yujia Zhou, Qingyi Pan, Blaze Chen, Yiqun Liu, Min Zhang, and Qingyao Ai. 2026. "EnGRICH: Enhancing Generative Reward Modeling with Critiques from Humans." https://omanscience.com/en/articles/engrich-enhancing-generative-reward-modeling-with-critiques-from-humans.

Harvard

Li, X., Wang, B., Li, H., Wang, H., Zhou, Y., Pan, Q., Chen, B., Liu, Y., Zhang, M. and Ai, Q. (2026) 'EnGRICH: Enhancing Generative Reward Modeling with Critiques from Humans', Available at: https://omanscience.com/en/articles/engrich-enhancing-generative-reward-modeling-with-critiques-from-humans.

Vancouver

Li X, Wang B, Li H, Wang H, Zhou Y, Pan Q, et al. EnGRICH: Enhancing Generative Reward Modeling with Critiques from Humans. https://omanscience.com/en/articles/engrich-enhancing-generative-reward-modeling-with-critiques-from-humans

IEEE

X. Li, B. Wang, H. Li, H. Wang, Y. Zhou, Q. Pan, B. Chen, Y. Liu, M. Zhang, and Q. Ai, "EnGRICH: Enhancing Generative Reward Modeling with Critiques from Humans," https://omanscience.com/en/articles/engrich-enhancing-generative-reward-modeling-with-critiques-from-humans.