Skip to content
AI Judge
Models
Bundles
Runs
Playground
Leaderboard
Compare
Judges
Settings
RUN #ed30
mini-benchmark-v1
CANCELLED
Show judge streams
5/40 tasks · 6:40 elapsed
Spend $0.1012 ($2.00 cap)
This run is not leaderboard-eligible: cancelled before completion.
Arena
Report
Model
Roleplay
Coding
Math
Research
Mktg
Poster
Story
Judging
avg
qwen3.8-max
qwen/qwen3.8-max
1.8
×5
—
—
—
—
—
—
1.8