Skip to content
AI Judge
Models
Bundles
Runs
Playground
Leaderboard
Compare
Judges
Settings
RUN #f9a0
mini-benchmark-v1
INCOMPLETE
Show judge streams
16/16 tasks · 15:01 elapsed
Spend $0.6948 ($2.00 cap)
Included on the leaderboard with penalties / reduced coverage. Infra failures score 0 (retry to replace); judging failures are excluded.
Arena
Report
Model
Roleplay
Coding
Math
Research
Mktg
Poster
Story
Judging
avg
laguna-s-2.1:free
poolside/laguna-s-2.1:free
✕
9.1
×2
6.1
×2
9.4
×2
✕
8.8
×2
7.3
×2
9.3
×2
8.3