Preprint Open access
CRJudgeBench: Can AI Detect Plausible but Invalid Code Reviews?
Large language models can generate plausible code-review comments, but such comments may contain technically incorrect claims that mislead developers. We study technical trustworthiness judgment: determining whether a review comment's core technical claims are correct and applicable to the reviewed code in its reposito …