How reliable are the judges themselves?
Agreement is a diagnostic, not a target — a well-supported minority judgment is not penalized. Scores here affect judge meta-rating only and never alter candidate rankings.
| Model | Judgments | Harshness | Variance | Parse fails | Evidence | Claim Δ | |
|---|---|---|---|---|---|---|---|
| claude-sonnet-4.6anthropic | 24 | +0.5 → lenient | 1.3Panel-wide σ 3.8 | 67% | 3.6 | 0.2 | |
| claude-sonnet-5anthropic | 58 | -0.4 harsh ← | 0.5Panel-wide σ 3.8 | 31% | 5.7 | 0.2 | |
| deepseek-v4-prodeepseek | 42 | -0.1 harsh ← | 0.7Panel-wide σ 3.8 | 50% | 4.8 | 0.5 | |
| gemini-3.1-pro-previewgoogle | 54 | +0.3 → lenient | 0.4Panel-wide σ 3.8 | 20% | 5.1 | 0.3 | |
| gemini-3.6-flashgoogle | 24 | +0.2 → lenient | 0.3Panel-wide σ 3.8 | 13% | 5.2 | 0.3 | |
| gpt-5.6-luna-proopenai | 38 | +1.2 → lenient | 0.9Panel-wide σ 3.8 | 3% | 5.6 | 1.9 | |
| gpt-5.6-solopenai | 37 | -0.1 harsh ← | 0.9Panel-wide σ 3.8 | 11% | 5.8 | 0.4 | |
| gpt-5.6-terraopenai | 45 | -0.2 harsh ← | 0.4Panel-wide σ 3.8 | 0% | 6.0 | 0.3 | |
| grok-4.5x-ai | 152 | -0.1 harsh ← | 0.6Panel-wide σ 3.8 | 1% | 6.3 | 0.5 | |
| deepseek-v4-flash-latest~deepseek | 6 | -0.6 harsh ← | 0.7Panel-wide σ 3.8 | 17% | 7.2 | 0.5 |
No calibration fixtures run yet. Fixtures are optional in v1.