Score: 0

Causal Judge Evaluation: Calibrated Surrogate Metrics for LLM Systems

Published: December 11, 2025 | arXiv ID: 2512.11150v1

By: Eddie Landesberg

LLM-as-judge evaluation has become the de facto standard for scaling model assessment, but the practice is statistically unsound: uncalibrated scores can invert preferences, naive confidence intervals on uncalibrated scores achieve near-0% coverage, and importance-weighted estimators collapse under limited overlap despite high effective sample size (ESS). We introduce Causal Judge Evaluation (CJE), a framework that fixes all three failures. On n=4,961 Chatbot Arena prompts (after filtering from 5k), CJE achieves 99% pairwise ranking accuracy at full sample size (94% averaged across configurations), matching oracle quality, at 14x lower cost (for ranking 5 policies) by calibrating a 16x cheaper judge on just 5% oracle labels (~250 labels). CJE combines three components: (i) AutoCal-R, reward calibration via mean-preserving isotonic regression; (ii) SIMCal-W, weight stabilization via stacking of S-monotone candidates; and (iii) Oracle-Uncertainty Aware (OUA) inference that propagates calibration uncertainty into confidence intervals. We formalize the Coverage-Limited Efficiency (CLE) diagnostic, which explains why IPS-style estimators fail even when ESS exceeds 90%: the logger rarely visits regions where target policies concentrate. Key findings: SNIPS inverts rankings even with reward calibration (38% pairwise, negative Kendall's tau) due to weight instability; calibrated IPS remains near-random (47%) despite weight stabilization, consistent with CLE; OUA improves coverage from near-0% to ~86% (Direct) and ~96% (stacked-DR), where naive intervals severely under-cover.

How to Correctly Report LLM-as-a-Judge Evaluations

Machine Learning (CS)

Fixes computer judge mistakes for fairer tests.

26 Nov 2025 1

88%

Aligning Evaluation with Clinical Priorities: Calibration, Label Shift, and Error Costs

Machine Learning (CS)

Helps doctors pick the best treatment for patients.

17 Jun 2025 0

88%

LLM-as-a-Judge for Scalable Test Coverage Evaluation: Accuracy, Operational Reliability, and Cost

Software Engineering

Tests computer programs faster and cheaper.

1 Dec 2025 3

View PDF Login to Bookmark

Causal Judge Evaluation: Calibrated Surrogate Metrics for LLM Systems

Technical Abstract

How to Correctly Report LLM-as-a-Judge Evaluations

Aligning Evaluation with Clinical Priorities: Calibration, Label Shift, and Error Costs

LLM-as-a-Judge for Scalable Test Coverage Evaluation: Accuracy, Operational Reliability, and Cost