Skip to content
AI Judge
Models
Bundles
Runs
Playground
Leaderboard
Compare
Judges
Settings
RUN #56c9
mini-benchmark-v1
INCOMPLETE
Show judge streams
0/8 tasks · 0:00 elapsed
Spend $0.0000 ($2.00 cap)
Included on the leaderboard with penalties / reduced coverage. Infra failures score 0 (retry to replace); judging failures are excluded.
Arena
Report
Model
Roleplay
Coding
Math
Research
Mktg
Poster
Story
Judging
avg
claude-sonnet-5
anthropic/claude-sonnet-5
—
—
—
—
—
—
—
—
—