How reliable are the judges themselves?
Agreement is a diagnostic, not a target — a well-supported minority judgment is not penalized. Scores here affect judge meta-rating only and never alter candidate rankings.
| Model | Judgments | Harshness | Variance | Parse fails | Evidence | Claim Δ | |
|---|---|---|---|---|---|---|---|
| claude-opus-4.5anthropic | 144 | -0.8 harsh ← | 0.6Panel-wide σ 0.7 | 2% | 9.1 | 0.3 | |
| o3openai | 144 | +0.6 → lenient | 0.7Panel-wide σ 0.7 | 4% | 8.7 | 0.5 | |
| gemini-3-progoogle | 96 | +0.2 → lenient | 0.8Panel-wide σ 0.7 | 6% | 8.2 | 0.7 | |
| deepseek-r1deepseek | 96 | -1.7outlier harsh ← | 1.1Panel-wide σ 0.7 | 14% | 6.8 | 1.6 | |
| kimi-k3moonshotai | 48 | +0.1 → lenient | 0.6Panel-wide σ 0.7 | 3% | 8.9 | 0.4 |
| Fixture | Judge | Evidence | Consistency | Correct | Parse |
|---|---|---|---|---|---|
| math-wrong-sum | claude-opus-4.5 | 8.6 | 8.9 | ✓ | first_try |
| poster-over-limit | claude-opus-4.5 | 8.6 | 8.9 | ✓ | first_try |
| coding-shape-only | claude-opus-4.5 | 8.6 | 8.9 | ✓ | first_try |
| story-perfect-wordcount | claude-opus-4.5 | 8.6 | 8.9 | ✓ | first_try |
| research-fabricated-citation | claude-opus-4.5 | 8.6 | 8.9 | ✓ | first_try |
| roleplay-missing-entries | claude-opus-4.5 | 8.6 | 8.9 | ✓ | first_try |