Abstract

An open contrastive decision model is near chance as a judge on the hard public benchmarks: Contrastive-LM/CLM-v0.1-8B scores between 0.351 (best- of-four, chance 0.250) and 0.593 (pairwise, chance 0.500), is statistically indistinguishable from coin flipping on RM-Bench and JudgeBench, and answers every HaluEval item with one constant label, matching the trivial always-first baseline at 0.581. Judges with the same parameter count score far higher everywhere: a reward model reaches 0.764 to 0.976 and a generative judge 0.611 to 0.778, and every gap to CLM is significant after Benjamini-Hochberg correction. Two properties do work. Raw confidences are overconfident by up to +0.401, yet one pooled temperature fit on held-out calibration items repairs expected calibration error to at most 0.062, and the repaired confidence ranks the model's own errors above chance on three of six benchmarks. The decision order-flip rate is 0.0002 against 0.2188 for the generative judge, and the length-preference shift is -0.023 against -0.217. The confidence-gated cascade, however, escalates between 0.923 and 1.000 of items to the strong judge at the preregistered 0.97 retention bar: calibrated confidence about a near-chance judge has almost nothing to keep. The design: five public preference benchmarks and one hallucination benchmark with real labels, scored under a preregistration frozen before any test item was seen, against generative, reward-model, and trivial baselines, with per-item predictions released.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Nandakishore, G. (2026). CLM-as-a-Judge: Evaluating an Open Contrastive Decision Model on Public Judge Benchmarks. https://omanscience.com/en/articles/clm-as-a-judge-evaluating-an-open-contrastive-decision-model-on-public-judge-benchmarks

MLA 9

Nandakishore, Gowthamkumar. "CLM-as-a-Judge: Evaluating an Open Contrastive Decision Model on Public Judge Benchmarks." https://omanscience.com/en/articles/clm-as-a-judge-evaluating-an-open-contrastive-decision-model-on-public-judge-benchmarks.

Chicago (author–date)

Nandakishore, Gowthamkumar. 2026. "CLM-as-a-Judge: Evaluating an Open Contrastive Decision Model on Public Judge Benchmarks." https://omanscience.com/en/articles/clm-as-a-judge-evaluating-an-open-contrastive-decision-model-on-public-judge-benchmarks.

Harvard

Nandakishore, G. (2026) 'CLM-as-a-Judge: Evaluating an Open Contrastive Decision Model on Public Judge Benchmarks', Available at: https://omanscience.com/en/articles/clm-as-a-judge-evaluating-an-open-contrastive-decision-model-on-public-judge-benchmarks.

Vancouver

Nandakishore G. CLM-as-a-Judge: Evaluating an Open Contrastive Decision Model on Public Judge Benchmarks. https://omanscience.com/en/articles/clm-as-a-judge-evaluating-an-open-contrastive-decision-model-on-public-judge-benchmarks

IEEE

G. Nandakishore, "CLM-as-a-Judge: Evaluating an Open Contrastive Decision Model on Public Judge Benchmarks," https://omanscience.com/en/articles/clm-as-a-judge-evaluating-an-open-contrastive-decision-model-on-public-judge-benchmarks.