About

SamePrompt gives every model the same prompt and shows what came back: the build, the cost, the time. Same prompt, different models — that is the whole method.

The problem

When someone says one model beats another, the prompt is in a video, the starting files are nowhere, and the measurement is in the author’s head. Nobody can check. This catalogue keeps the prompt in plain text next to what each model produced, so a claim can at least be read, and at best replayed.

Two kinds of entry

A replayable pack ships everything needed to run the test again: the verbatim prompt, any starting files, the measurement protocol, and a packSha256 identifying the exact material. Numbers attached to it were measured locally in Bench Arena, with tokens reconciled on the harness log and cost recomputed from them.

A referenced entry is a comparison published elsewhere. We link the source, credit the author, and mark every figure as reported. We did not run it and cannot verify it.

The two never share a number. There is no score, no average, no leaderboard: mixing a reconciled measurement with a figure from a screenshot would produce a ranking that looks rigorous and is not.

How numbers are shown

A value that was not measured reads not measured, never $0.00. A zero looks like a real, excellent result, and a wrong metric costs more than a missing one.

Every figure carries the outcome of its run. A low cost on a run that timed out is not a performance, and hiding the status would let it read as one.

Credit and removal

Clips are short, muted excerpts re-hosted so the wall stays fast. Each one names its author and links to the original post. Handles are read off the source, never reconstructed from a name: a guessed handle credits a stranger for someone else’s work.

If your work is here and you want it changed or removed, ask and it comes down within 48 hours. No justification needed.

Bench Arena

This site never runs an agent. Bench Arena is the companion local application: it runs on your machine, with your own subscriptions, and produces the measurements SamePrompt publishes. That split is deliberate — a single serious run can cost tens of dollars, so hosted runs for everyone was never a viable shape for this.

Browse the catalogue →