Four findings organize the study
Every judge configuration is compared under the same decision-only contract: a discrete judgment plus label probabilities, no rationale. Frozen inputs, retained failures, and source-cluster bootstrap intervals throughout.
Finding 01Where JEV is enough
Ties the strongest judge on preference and evidence-grounded factuality
RewardBench 92.2 vs 93.5 and HaluEval 87.5 vs 86.7 against GPT-6, both within the confidence interval on benchmark labels, with valid output on all 1,312 items. A blinded human adjudication of every item where the two judges differ puts GPT-6 two to three points ahead on these tasks — and finds that 24 of the 26 HaluEval items both judges “miss” carry labels the evidence does not support. Among hosted judges under $1 per 1,000 judgments, JEV leads on JudgeBench, ties on HaluEval, and sits 0.6 points off the best on RewardBench.
Finding 02When to escalate
Hard correctness and misleading style
JudgeBench: 78.6 vs 93.1, a 14.6-point gap concentrated in reasoning and coding — and it is not label noise: blinded human adjudication sides with GPT-6 on 57 of the 69 disputed items and with JEV on one (human-adjudicated gap −16.0, interval −20.0 to −12.3). RM-Bench: when the rejected answer is written in the more elaborate style, JEV loses 9.2 points and GPT-6 loses none. Selecting the right answer and resisting style are different abilities.
JudgeBench · RM-Bench hard
Finding 03A judge boundary
Reference-free prose: near chance, fully confident — for every judge
On eighty balanced reference-free responses, JEV, GPT-4.1 mini, and GPT-5.4 agree with only 52.5–55.0% of existing hallucination labels while reporting mean maximum probabilities of 0.90–0.96. Do not rely on any tested judge here.
Finding 04Confidence as a signal
An escalation signal, not a certificate
A frozen JEV → GPT-6 cascade keeps 99% of GPT-6's accuracy at 57% of its fee. But native probabilities are no more reliable than verbalized ones, and fitted temperatures worsened NLL on two of three benchmark families — validate per workload.