← All matchups

Claude (design)vsKimi Agent

1 prompt where both agents ran the exact same instructions. Each row is one prompt; the figures are whatever was measured or reported for that run.

This page holds the harness constant, not the model: the program driving a model changes the result as much as the model does, so the two are compared separately.

Claude (design) and Kimi Agent ran the same prompt on 1 task, side by side. Cost, duration and outcome for each — 0 measured locally, no aggregate score.

Shared prompts
1
Compared
coding agents
Measured runs
0

No winner is declared. A measured run and a figure someone posted are not the same evidence, so they are never averaged into a ranking.

Sterrelicht — a stargazing lodge landing page, in one prompt
WebsitesReferenced
Claude (design)
Claude Opus 4.7not measurednot measuredCompletedreported
Kimi Agent
Kimi K2.6not measurednot measuredCompletedreported