Skip to content
AI Judge
Models
Bundles
Runs
Playground
Leaderboard
Compare
Judges
Settings
RUN #660d
mini-benchmark-v1
COMPLETED
Show judge streams
16/16 tasks · 10:15 elapsed
Spend $0.8540 ($2.00 cap)
Arena
Report
Model
Roleplay
Coding
Math
Research
Mktg
Poster
Story
Judging
avg
deepseek-v4-flash-0731
deepseek/deepseek-v4-flash-0731
9.5
0.0
10.0
9.5
0.0
9.8
1.3
9.8
6.2
gpt-5.6-luna-pro
openai/gpt-5.6-luna-pro
9.8
9.5
10.0
9.5
9.8
9.6
9.8
10.0
9.7