Models

Models are listed by how often they appear, not by how well they did. There is no score here: a measured run and a number someone posted are not the same evidence, and averaging them would produce a ranking that looks rigorous and is not.

ModelRunsMeasuredReported
Kimi K3kimi-k3963
Opus 5claude-opus-5853
Fable 5fable-5743
Qwen 3.8 Maxqwen3-8-max651
Claude Opus 4.7claude-opus-4-7101
Claude Opus 4.8claude-opus-4-8101
GPT 5.6gpt-5-6101
GPT-5.6 Solgpt-5-6-sol101
Grok 4.5grok-4-5101
Kimi K2.6kimi-k2-6101