← All matchups

Opus 5vsQwen 3.8 Max

3 prompts where both models ran the exact same instructions. Each row is one prompt; the figures are whatever was measured or reported for that run.

Opus 5 and Qwen 3.8 Max ran the same prompt on 3 tasks, side by side. Cost, duration and outcome for each — 7 measured locally, no aggregate score.

Shared prompts
3
Compared
models
Measured runs
7

No winner is declared. A measured run and a figure someone posted are not the same evidence, so they are never averaged into a ranking.

Halo — Blood Gulch 5v5 in the browser
GamesReferenced
Opus 5
not measured12h 00mCompletedreported
Qwen 3.8 Max
not measurednot measuredCompletedreported
Jelly Jungle — 3D jungle run
GamesReplayable pack
Opus 5
$20.92h 00mTimed outmeasured
$21.31h 09mFailedmeasured
Qwen 3.8 Max
$1.0222m 56sCompletedmeasured
$1.2734m 46sCompletedmeasured
Wobble Rush — 3D obstacle course
GamesReplayable pack
Opus 5
$9.3923m 15sInterruptedmeasured
$10.534m 41sCompletedmeasured
Qwen 3.8 Max
$1.5830m 46sCompletedmeasured

What this page shows, and how to read it

Opus 5 is a large language model from Anthropic. On the prompts below, it ran through Codex CLI and Claude Code. Qwen 3.8 Max is a large language model from Alibaba. On the prompts below, it ran through OpenCode and Qwen Code.

The two sides share 3 prompts on this site: "Halo — Blood Gulch 5v5 in the browser", "Jelly Jungle — 3D jungle run" and "Wobble Rush — 3D obstacle course". Across them, Opus 5 completed 2 of 5 runs — of the others, 1 failed, 1 timed out and 1 interrupted, and Qwen 3.8 Max completed all 4 of its runs. Opus 5 has measured costs from $9.39 to $21.3, measured durations from 23m 15s to 2h 00m and a reported duration of 12h 00m. Qwen 3.8 Max has measured costs from $1.02 to $1.58 and measured durations from 22m 56s to 34m 46s.

How to read the outcomes. Completed means the prompt alone produced a working artefact. Failed means this run did not — the build broke, the result would not run, or the agent stopped short of a working state. That is a fact about one run on one prompt, not a verdict on the tool: the same programs complete other prompts elsewhere on this site. Timed out means the run was cut at its time limit with the artefact unfinished, so whatever it cost bought a partial result. Stopped and Interrupted mark runs ended from the outside before they concluded. A low cost attached to a run that did not finish is not a saving — it is the price of an attempt, which is why every figure on this page travels with its status.

Every figure above carries a trust label. Measured means SamePrompt ran it locally in Bench Arena, reconciled the tokens on the harness log and recomputed the cost from them. Reported means the author of the source announced the figure; it is shown as stated and cannot be verified here. The two are never summed, averaged or ranked, and this page declares no winner: a cheap run that failed and an expensive run that completed are two facts, not a score.