mini-benchmark-v1·hash a3f2c1d49b8e
| Model | Disagreement | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| 1 | claude-sonnet-4.5anthropic | 9.0 | 4 | 98% | $0.3100 | |||||
| 2 | gpt-5.1openai | 8.5 | 3 | 96% | $0.4200 | |||||
| 3 | deepseek-v4deepseek | 7.5 | 3 | 91% | $0.0200 | |||||
| — | gemini-3-progooglePROVISIONALfewer than 3 complete runs — median of 1 run shown | 8.8 | 1 | 100% | $0.2800 | |||||
| — | grok-4.1x-aiPROVISIONALfewer than 3 complete runs — median of 2 runs shown | 7.9 | 2 | 94% | $0.3300 |