Benchmark lab
AI Judge
One bundle. Three independent judges. Reproducible rankings.
A single-operator instrument for honest model comparison — streamed answers, deterministic validators, seeded blind judge panels, durable SQLite records.
Methodology
Bundle
One immutable, versioned prompt bundle — 8 category tasks under a common wrapper. Same input for every model, forever.
Stream
Candidates answer live over OpenRouter. Deterministic validators check JSON, counts, word limits and known answers first.
Judge ×3
A seeded blind panel of three judges scores each answer at temperature 0. Candidate identity is never revealed.
Rank
Median of judge overalls per task, macro-averaged across categories. Only complete runs enter the leaderboard.
Current standings
mini-benchmark-v1
| Rank | Model | Median | Runs | Last evaluated |
|---|---|---|---|---|
| — | openai/gpt-5.6-lunaPROVISIONAL | 9.8 | 1 | 26d ago |
| — | openai/gpt-5.6-luna-proPROVISIONAL | 9.2 | 2 | 26d ago |
| — | x-ai/grok-4.5PROVISIONAL | 8.7 | 1 | 26d ago |
| — | anthropic/claude-sonnet-4.6PROVISIONAL | 8.3 | 1 | 26d ago |
| — | deepseek/deepseek-v4-flashPROVISIONAL | 7.8 | 1 | 26d ago |
Why it is honest
Blind judging
Judge prompts contain only the task and the raw answer — never the model's name, provider, or metadata.
Seeded panels
Each category gets one deterministic 3-judge panel per run, persisted with its seed and reserve order. Fully reproducible.
Deterministic validators
Objective checks — schema, counts, word limits, known math answers — run before judging and are shown separately, never blended away.