Side-by-side scores for up to four models on one bundle version.
| Category | gpt-5.6-luna | gpt-5.6-luna-pro |
|---|---|---|
| Roleplay | 10.0±0.3 panel low: 9.9median: 10.0panel high: 10.0 | 9.8±0.4 panel low: 9.5median: 9.8panel high: 10.0 |
| Coding | 9.6±0.5 panel low: 9.4median: 9.6panel high: 9.9 | 8.0±0.4 panel low: 7.8median: 8.0panel high: 8.2 |
| Math | 10.0±0.0 panel low: 10.0median: 10.0panel high: 10.0 | 10.0±0.1 panel low: 10.0median: 10.0panel high: 10.0 |
| Research | 9.8±0.5 panel low: 9.5median: 9.8panel high: 10.0 | 9.6±0.6 panel low: 9.3median: 9.6panel high: 9.9 |
| Marketing | 9.9±1.0 panel low: 9.4median: 9.9panel high: 10.0 | 8.2±0.1 panel low: 8.2median: 8.2panel high: 8.3 |
| Poster | 9.5±0.5 panel low: 9.3median: 9.5panel high: 9.8 | 8.9±0.6 panel low: 8.6median: 8.9panel high: 9.2 |
| Story | 9.8±0.5 panel low: 9.6median: 9.8panel high: 10.0 | 9.0±0.4 panel low: 8.7median: 9.0panel high: 9.2 |
| Judging | 9.8±0.8 panel low: 9.4median: 9.8panel high: 10.0 | 10.0±0.5 panel low: 9.8median: 10.0panel high: 10.0 |
db053025 · 26d ago
Best improvement: Provide a small self-contained test harness defining `assert` and `assertRejects` so the listed tests can be run directly without relying on unspecified globals.
db053025 · 26d ago
Best improvement: Replace timer-based eviction with lazy expiration checks, or schedule long TTLs in bounded chunks, and provide executable TypeScript tests with concrete assertions rather than prose descriptions.
| Model | Runs | Incomplete | Success | Median / IQR | Cost / run | Latency | Score / $ |
|---|---|---|---|---|---|---|---|
| gpt-5.6-luna | 1 | 0 | 100% | 9.8(9.8–9.8) | $0.0044 | 7.7s | 2233.7 pts/$ |
| gpt-5.6-luna-pro | 2 | 0 | 100% | 9.2(8.9–9.5) | $0.0194 | 16.3s | 474.3 pts/$ |