Preprint · September 2026

Accept when confident, escalate when unsure.

JEV-as-a-Judge measures a typed decision service — a judge that returns only a verdict and a probability distribution — against sixteen generative and reward-model configurations. On ordinary preference and evidence-grounded factuality it stays within three points of GPT-6 at 0.36% of the fee, with valid output on every item; the same measurements map exactly where a stronger judge is still worth paying for.

Yubo Li, Yidi Miao, Ramayya Krishnan, Rema Padman
Carnegie Mellon University

0.152s
Median latency on a matched 120-judgment panel
$0.044
Per 1,000 judgments, estimated from reported usage
277×
Cheaper than GPT-6 at low effort; 9× cheaper than GPT-4.1 mini
14.6pt
Behind GPT-6 on JudgeBench, while within 1.3 points on RewardBench

All numbers are from the manuscript (September 2026). Fees are estimates from reported token usage at collection-time prices, not invoices; configurations differ in model size and reasoning effort and are not compute-matched.

Four findings organize the study

Every judge configuration is compared under the same decision-only contract: a discrete judgment plus label probabilities, no rationale. Frozen inputs, retained failures, and source-cluster bootstrap intervals throughout.

Finding 01Where JEV is enough

Ties the strongest judge on preference and evidence-grounded factuality

RewardBench 92.2 vs 93.5 and HaluEval 87.5 vs 86.7 against GPT-6, both within the confidence interval on benchmark labels, with valid output on all 1,312 items. A blinded human adjudication of every item where the two judges differ puts GPT-6 two to three points ahead on these tasks — and finds that 24 of the 26 HaluEval items both judges “miss” carry labels the evidence does not support. Among hosted judges under $1 per 1,000 judgments, JEV leads on JudgeBench, ties on HaluEval, and sits 0.6 points off the best on RewardBench.

Finding 02When to escalate

Hard correctness and misleading style

JudgeBench: 78.6 vs 93.1, a 14.6-point gap concentrated in reasoning and coding — and it is not label noise: blinded human adjudication sides with GPT-6 on 57 of the 69 disputed items and with JEV on one (human-adjudicated gap −16.0, interval −20.0 to −12.3). RM-Bench: when the rejected answer is written in the more elaborate style, JEV loses 9.2 points and GPT-6 loses none. Selecting the right answer and resisting style are different abilities.

Finding 03A judge boundary

Reference-free prose: near chance, fully confident — for every judge

On eighty balanced reference-free responses, JEV, GPT-4.1 mini, and GPT-5.4 agree with only 52.5–55.0% of existing hallucination labels while reporting mean maximum probabilities of 0.90–0.96. Do not rely on any tested judge here.

Finding 04Confidence as a signal

An escalation signal, not a certificate

A frozen JEV → GPT-6 cascade keeps 99% of GPT-6's accuracy at 57% of its fee. But native probabilities are no more reliable than verbalized ones, and fitted temperatures worsened NLL on two of three benchmark families — validate per workload.

Task fit matters more than a universal ranking

Seventeen configurations, base-order accuracy, invalid outputs counted as errors. Hosted reasoning models run at low reasoning effort (Qwen3.6 at its default); local Qwen uses non-thinking inference. Switch the task: the order re-sorts. JEV is within 1.3 points of GPT-6 on RewardBench and 14.6 points behind on JudgeBench (95% source-cluster interval −18.9 to −10.3); under blinded human adjudication of the disputed items the gaps are 3.0 and 16.0 points.

JEV 1.13GPT-6 AstraReward modelsOther generative judges

View data

Paired JEV − GPT-6 difference

Percentage points, 95% source-cluster bootstrap interval, 2,000 resamples.

Near-equal RewardBench and HaluEval scores do not establish general non-inferiority: the intervals allow meaningful differences. JudgeBench separates the two judges cleanly.

Where the JudgeBench gap lives

Accuracy by domain, JEV vs GPT-6 (n per domain).

Reasoning and coding drive the gap; knowledge differs least. Solving difficult comparisons is a distinct capability from ordinary preference judging.

Measured latency and fees

A separate frozen 120-decision panel, one model group at a time, eight workers, 0.12-second minimum starts. Outcome latency includes network, provider queueing, and retries after the first attempt. Fees use reported usage at collection-time prices; missing usage keeps a conservative reservation.

Outcome latency

Median (dot) and p95 (bar end), seconds, log scale.

Estimated fee per 1,000 judgments

Reported usage (dot); lighter extension = conservative reservation where usage was missing. USD, log scale.

The roughly 9-fold and 277-fold fee differences describe this workload and these services. Different model capacities, reasoning budgets, caches, and providers prevent attributing the gap solely to a non-generative architecture.

Harder selection and answer style

Follow-up samples that RewardBench 1 cannot expose: 100 RewardBench 2 prompts (four-way top-1 selection) and 80 RM-Bench prompts with all nine style pairings in both orders. JEV loses 9.2 points when the rejected answer is written in the more elaborate style; GPT-6 does not.

RM-Bench by style condition

Accuracy (%), 480 judgments per condition. Hard: the less elaborate answer is the preferred one.

JEVGPT-6 AstraSkywork-Reward-V2 8B

Hard-minus-normal change: JEV −9.2 [−14.0, −4.8], GPT-6 +1.3 [−1.9, 4.2], Skywork −20.6 [−26.9, −14.4]. On RewardBench 2 four-way selection the three score 73.0 / 75.0 / 79.0.

All nine style combinations

Rows: preferred answer's style. Columns: rejected answer's style. Each cell: 160 judgments.

Above the diagonal the rejected answer is more elaborate. JEV's top-right cell (concise preferred vs. Markdown rejected) falls to 70.6%; GPT-6 stays above 95%.

Confidence is not a correctness certificate

JEV's native confidence is a statistic of its returned distribution (Spearman 0.95–1.00 with the maximum label probability), so the paper uses max probability q for every judge. Three views of the same lesson.

Reference-free hallucination judgments

Bar: label agreement (%). Tick: mean maximum probability. 80 balanced general responses, no evidence supplied.

All three judges are near chance and near-certain: JEV's Brier score is 0.815 with error-detection AUROC 0.518; GPT-5.4's Brier score is 0.813. On document-grounded summaries the same judges reach 71.2 / 62.5 / 72.5.

Accuracy and uncertainty rank differently

Brier score (lower is better), valid outputs only. Error-detection AUROC below.

GPT-6 improves both accuracy and probability quality on hard comparisons; its greater JudgeBench capability does not carry over to evidence-grounded uncertainty on HaluEval, where JEV's Brier score is better (0.176 vs 0.245) yet GPT-6's error ranking is stronger (AUROC 0.899 vs 0.863).

Temperature scaling does not transfer within a benchmark family

One temperature per task, fitted on the pilot selection set (64 / 32 / 32 judgments), evaluated on the disjoint extension. NLL, lower is better.

RewardBench T = 0.65 (sharpens)
0.2050.232
worse on 240 extension pairs
JudgeBench T = 2.15 (softens)
0.4490.475
worse on 270 extension pairs
HaluEval T = 4.45 (softens)
0.3330.284
better on 160 extension decisions

The fitted temperatures differ seven-fold across tasks. Task-specific validation beats a universal temperature.

3.2% / 11.1%
Reversal disagreement, RewardBench / JudgeBench (GPT-6: 1.5% / 0.9%)
0/96
Decision changes across repeated identical requests; 4/48 under a paraphrased rubric
48.9%
First-position rate across both orders — no aggregate position preference, but 39 inconsistent JudgeBench pairs
10/199
Errors at q ≥ 0.9 on HaluEval (GPT-6: 28/233); on JudgeBench, 9 of 138 high-probability judgments are wrong

When does escalation pay?

JEV's accuracy rises monotonically with its confidence — pooled over the three public tasks, from 48% below q = 0.6 to 99% at q = 1 — and its gap to GPT-6 sits almost entirely in the low-confidence decisions: on the items JEV accepts at q ≥ 0.9 the two judges are within a point of each other, on the rest GPT-6 leads by 15. So: accept when confident, escalate the rest. A frozen two-order gate does exactly that — JEV judges each pair in both orders, averages the aligned probability, and keeps the decision when it is confident; everything else goes to a fallback judge. Thresholds were selected on 96 pilot pairs and evaluated on 510 disjoint extension pairs. The signal weakens where JEV is confidently misled: on RM-Bench's style-adversarial pairs its error-detection AUROC drops from 0.90 to 0.77, and the cascade needs a higher threshold to keep pace.

JEV → GPT-6 at threshold 0.9

510 extension preference pairs (RewardBench + JudgeBench)two JEV calls per pair, fallback only when escalated
92.5%vs 93.1% GPT-6 alone
Cascade accuracy. Paired change −0.59 points, 95% interval [−1.78, 0.59].
43%fee reduction
Uses 56.8% of the estimated GPT-6 fee (conservative upper ratio 62.2%), including both JEV orders.
60of 75
JudgeBench errors of JEV that GPT-6 corrects; JEV corrects 9 of GPT-6's 24. Oracle union 95.7% — an upper bound that needs labels.
Other fallbacks
Fallback judgeτAccepted by JEVCascade acc.Fallback aloneΔ (pp) [95% CI]Fee ratio (reported / upper)
GPT-5.40.9053.7%91.491.6−0.2 [−1.2, 0.8]0.639 / 1.096
GPT-5.6 Sol0.7081.0%91.093.3−2.4 [−4.5, −0.2]0.288 / 0.380
GPT-6 Astra0.9053.7%92.593.1−0.6 [−1.8, 0.6]0.568 / 0.622
Claude Sonnet 50.6090.8%88.089.2−1.2 [−4.0, 1.8]0.152 / 0.152
Gemini 3 Flash0.6090.8%87.182.9+4.1 [1.2, 7.2]0.281 / 0.282
Gemini 3.1 Pro0.9053.7%89.690.4−0.8 [−1.8, 0.2]0.594 / 0.635

The GPT-5.6 policy accepts 81% of pairs but loses 2.35 points on transfer — more than the two-point tolerance it was selected under. Tolerances must be re-checked on the target workload; these are offline simulations, not live cascade latency.

Usable output is a distinct endpoint

Primary accuracy counts every failed call as an error. Provider-side structured-generation failures, HTTP-200 contract violations, and transport failures are separated so that service availability is not mistaken for reasoning quality.

Hosted judgeValid / 1,312GenerationContractTransportBreakdownJudgeBench acc. (all / valid only)

Qwen3.6 at default effort would be a strong judge on its valid outputs (87.5% on JudgeBench) but loses 106 of 1,312 items to rate limits and malformed JSON — a deployment endpoint separate from conditional accuracy.

Reproducibility

Frozen requests, retained failed attempts, complete probability traces, a separate timing panel, and pre-specified routing thresholds. The anonymous supplement reproduces every public-task number offline — no API calls, no GPU. Prompts, schemas, and an exact JEV request/response are in the paper's appendix.

Not compute-matched. One proprietary JEV version. Benchmark agreement, not adjudicated correctness. Estimates, not invoices. The paper's Limitations section is the honest reading guide.

BibTeX

@misc{li2026jevjudge,
  title   = {{JEV}-as-a-Judge: Accept When Confident,
             Escalate When Unsure},
  author  = {Li, Yubo and Miao, Yidi and Krishnan, Ramayya
             and Padman, Rema},
  year    = {2026},
  note    = {Preprint}
}