Models
Models are listed by how often they appear, not by how well they did. There is no score here: a measured run and a number someone posted are not the same evidence, and averaging them would produce a ranking that looks rigorous and is not.
| Model | Runs | Measured | Reported |
|---|---|---|---|
| Opus 5claude-opus-5 | 14 | 5 | 9 |
| Kimi K3 (high)kimi-k3 | 13 | 7 | 6 |
| Fable 5fable-5 | 9 | 5 | 4 |
| Qwen 3.8 Maxqwen3-8-max | 6 | 5 | 1 |
| GPT-5.6gpt-5-6 | 2 | 0 | 2 |
| Grok 4.5 (high)grok-4-5 | 2 | 0 | 2 |
| 1-bit Kimi K3kimi-k3-1bit | 1 | 0 | 1 |
| Claude Opus 4.7claude-opus-4-7 | 1 | 0 | 1 |
| Claude Opus 4.8claude-opus-4-8 | 1 | 0 | 1 |
| DeepSeek V4 Flash-0731deepseek-v4-flash-0731 | 1 | 0 | 1 |
| GLM 5.2 (high)glm-5-2 | 1 | 0 | 1 |
| GPT-5.6 Solgpt-5-6-sol | 1 | 0 | 1 |
| Kimi K2.6kimi-k2-6 | 1 | 0 | 1 |
| Luna 5.6 (max reasoning)luna-5-6 | 1 | 0 | 1 |