الملخص
Peer-review evaluation is increasingly being automated with LLM-as-a-judge metrics, but this creates a measurement risk. A review may receive a high score because it is fluent, organized, and polished, rather than because it provides a strong evaluation of the paper. This risk is especially important in AI-assisted reviewing, where reviewers may use LLMs to improve clarity or presentation while preserving the underlying judgments. We propose a statistical framework for testing whether peer-review evaluation metrics capture substantive review quality beyond surface-level linguistic form. The framework compares original human reviews with faithful LLM rewrites that preserve the same evaluative content while changing wording and presentation. Using a dataset comprising 4,044 meaning-preserving rewrites derived from 674 human reviews from ICLR and NeurIPS, we evaluate 29 content-oriented peer-review evaluation metrics drawn from four prior works through complementary tests of surface sensitivity and robustness. Although these metrics are intended to capture review properties beyond surface-level, writing-dependent characteristics, we find that sensitivity to rewriting is widespread. Under our primary analysis, 23 metrics assign significantly different scores to reviews whose evaluative content is preserved, while only six satisfy our robustness criterion. The patterns are largely consistent across two LLM judge models, suggesting that the issue is not specific to a single judge. These findings show that many peer-review evaluation metrics partially conflate review quality with linguistic presentation, and indicate that robustness to meaning-preserving rewriting should be validated before such metrics are used to compare human-written, AI-assisted, and AI-generated reviews.
الكلمات المفتاحية
الموضوع
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Amirshahi, S., Ebrahimi, S., Le, H. S., Arabzadeh, N., & Bagheri, E. (2026). Judging a Review by its Cover: A Reliability Analysis of LLM-based Peer Review Evaluation Metrics. https://omanscience.com/ar/articles/judging-a-review-by-its-cover-a-reliability-analysis-of-llm-based-peer-review-evaluation-metrics
MLA 9
Amirshahi, Shakiba, et al. "Judging a Review by its Cover: A Reliability Analysis of LLM-based Peer Review Evaluation Metrics." https://omanscience.com/ar/articles/judging-a-review-by-its-cover-a-reliability-analysis-of-llm-based-peer-review-evaluation-metrics.
شيكاغو (المؤلف–التاريخ)
Amirshahi, Shakiba, Sajad Ebrahimi, Hai Son Le, Negar Arabzadeh, and Ebrahim Bagheri. 2026. "Judging a Review by its Cover: A Reliability Analysis of LLM-based Peer Review Evaluation Metrics." https://omanscience.com/ar/articles/judging-a-review-by-its-cover-a-reliability-analysis-of-llm-based-peer-review-evaluation-metrics.
هارفارد
Amirshahi, S., Ebrahimi, S., Le, H. S., Arabzadeh, N. and Bagheri, E. (2026) 'Judging a Review by its Cover: A Reliability Analysis of LLM-based Peer Review Evaluation Metrics', Available at: https://omanscience.com/ar/articles/judging-a-review-by-its-cover-a-reliability-analysis-of-llm-based-peer-review-evaluation-metrics.
فانكوفر
Amirshahi S, Ebrahimi S, Le HS, Arabzadeh N, Bagheri E. Judging a Review by its Cover: A Reliability Analysis of LLM-based Peer Review Evaluation Metrics. https://omanscience.com/ar/articles/judging-a-review-by-its-cover-a-reliability-analysis-of-llm-based-peer-review-evaluation-metrics
IEEE
S. Amirshahi, S. Ebrahimi, H. S. Le, N. Arabzadeh, and E. Bagheri, "Judging a Review by its Cover: A Reliability Analysis of LLM-based Peer Review Evaluation Metrics," https://omanscience.com/ar/articles/judging-a-review-by-its-cover-a-reliability-analysis-of-llm-based-peer-review-evaluation-metrics.