AI Judge

Benchmark lab

AI Judge

One bundle. Three independent judges. Reproducible rankings.

A single-operator instrument for honest model comparison — streamed answers, deterministic validators, seeded blind judge panels, durable SQLite records.

Methodology

01

Bundle

One immutable, versioned prompt bundle — 8 category tasks under a common wrapper. Same input for every model, forever.

02

Stream

Candidates answer live over OpenRouter. Deterministic validators check JSON, counts, word limits and known answers first.

03

Judge ×3

A seeded blind panel of three judges scores each answer at temperature 0. Candidate identity is never revealed.

04

Rank

Median of judge overalls per task, macro-averaged across categories. Only complete runs enter the leaderboard.

Current standings

mini-benchmark-v1

Full leaderboard →
Top models on mini-benchmark-v1
RankModelMedian
openai/gpt-5.6-lunaPROVISIONAL9.8
openai/gpt-5.6-luna-proPROVISIONAL9.2
x-ai/grok-4.5PROVISIONAL8.7
anthropic/claude-sonnet-4.6PROVISIONAL8.3
deepseek/deepseek-v4-flashPROVISIONAL7.8
Median of complete bundle-run scores · provisional < 3 complete runs · score 9.8 leader

Why it is honest

Blind judging

Judge prompts contain only the task and the raw answer — never the model's name, provider, or metadata.

Seeded panels

Each category gets one deterministic 3-judge panel per run, persisted with its seed and reserve order. Fully reproducible.

Deterministic validators

Objective checks — schema, counts, word limits, known math answers — run before judging and are shown separately, never blended away.

mini-benchmark-v1 · hash 825e262c4771Single-operator benchmark lab · SQLite WAL · temperature-0 judging