Skip to content
AI Judge
Models
Bundles
Runs
Playground
Leaderboard
Compare
Judges
Settings
RUN #9ddd
mini-benchmark-v1
COMPLETED
Show judge streams
8/8 tasks · 10:00 elapsed
Spend $0.5331 ($2.00 cap)
Arena
Report
Model
Roleplay
Coding
Math
Research
Mktg
Poster
Story
Judging
avg
claude-opus-5
anthropic/claude-opus-5
10.0
0.0
10.0
3.8
9.8
9.8
4.5
10.0
7.2