mini-benchmark-v1·hash 825e262c4771
| Model | Disagreement | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| — | gpt-5.6-lunaopenaiPROVISIONALfewer than 3 complete runs — median of 1 run shown | 9.8 | 1 | 100% | $0.0044 | |||||
| — | gpt-5.6-luna-proopenaiPROVISIONALfewer than 3 complete runs — median of 2 runs shown | 9.2 | 2 | 100% | $0.0194 | |||||
| — | grok-4.5x-aiPROVISIONALfewer than 3 complete runs — median of 1 run shown | 8.7 | 1 | 100% | $0.0692 | |||||
| — | claude-sonnet-4.6anthropicPROVISIONALfewer than 3 complete runs — median of 1 run shown | 8.3 | 1 | 100% | $0.1153 | |||||
| — | deepseek-v4-flashdeepseekPROVISIONALfewer than 3 complete runs — median of 1 run shown | 7.8 | 1 | 100% | $0.0017 | |||||
| — | claude-sonnet-5anthropicPROVISIONALfewer than 3 complete runs — median of 1 run shown | 7.4 | 1 | 100% | $0.1110 | |||||
| — | claude-opus-5anthropicPROVISIONALfewer than 3 complete runs — median of 2 runs shown | 6.9 | 2 | 100% | $0.3156 | |||||
| — | nemotron-3-ultra-550b-a55b:freenvidiaPROVISIONALfewer than 3 complete runs — median of 2 runs shown | 6.8 | 2 | 100% | $0.0000 | |||||
| — | deepseek-v4-flash-0731deepseekPROVISIONALfewer than 3 complete runs — median of 2 runs shown | 6.6 | 2 | 100% | $0.0030 | |||||
| — | laguna-s-2.1:freepoolsidePROVISIONALfewer than 3 complete runs — median of 1 run shown | 4.5 | 1 | 63% | $0.0000 | |||||
| — | qwen3.7-plusqwenPROVISIONALfewer than 3 complete runs — median of 1 run shown | 3.6 | 1 | 100% | $0.0180 | |||||
| — | qwen3.8-maxqwenPROVISIONALfewer than 3 complete runs — median of 1 run shown | 2.4 | 1 | 100% | $0.0891 | |||||
| — | qwen3.7-maxqwenPROVISIONALfewer than 3 complete runs — median of 1 run shown | 2.2 | 1 | 100% | $0.0669 | |||||
| — | qwen3.6-plusqwenPROVISIONALfewer than 3 complete runs — median of 1 run shown | 0.7 | 1 | 100% | $0.0293 |