نسخة أولية وصول مفتوح
Judging a Review by its Cover: A Reliability Analysis of LLM-based Peer Review Evaluation Metrics
Peer-review evaluation is increasingly being automated with LLM-as-a-judge metrics, but this creates a measurement risk. A review may receive a high score because it is fluent, organized, and polished, rather than because it provides a strong evaluation of the paper. This risk is especially important in AI-assisted rev …