Skip to content
AI Judge
Models
Bundles
Runs
Playground
Leaderboard
Compare
Judges
Settings
RUN #26e3
mini-benchmark-v1
COMPLETED
Show judge streams
32/32 tasks · 28:21 elapsed
Spend $0.9604 ($2.00 cap)
Arena
Report
Model
Roleplay
Coding
Math
Research
Mktg
Poster
Story
Judging
avg
qwen3.6-plus
qwen/qwen3.6-plus
1.3
1.3
0.0
0.0
1.3
0.0
1.3
0.5
0.7
qwen3.7-max
qwen/qwen3.7-max
1.3
2.8
10.0
1.3
0.0
0.5
1.3
0.5
2.2
qwen3.7-plus
qwen/qwen3.7-plus
1.3
3.5
10.0
1.3
1.3
1.3
1.3
9.4
3.6
qwen3.8-max
qwen/qwen3.8-max
0.0
0.0
5.0
9.5
0.0
1.3
0.3
3.0
2.4