OpenRouter provides the cost and speed data. The rest are benchmarks, weighted by the workload you picked to arrive at quality.
Every score on this page is somebody else’s measurement, carried over unchanged with the date it was taken.
Nothing here is scored by me, and nothing is a vendor’s own claim. Against each source is what it can tell you, what it cannot, and how much of this catalog it actually reaches.
Tells you: How reliably a model uses tools across single- and multi-turn tasks. V4 also tests search and memory. Native function-calling results are preferred; prompt-based results are used when those are the only measurements.
Does not tell you: Performance under every agent setup or on the newest models. The tested mode stays with the score; the date shown is the leaderboard update date, not an individual run date.
Tells you: One team running every model under one protocol, rather than collected vendor claims: GPQA Diamond, FrontierMath, SimpleQA Verified, SWE-Bench Verified and AIME.
Does not tell you: Anything about your domain. These are exam questions with one right answer; most work is not. AIME is kept on the page but not ranked on — the frontier scores 1.000, so it no longer separates two good models.
Tells you: Conversation skills and creative writing quality.EQ-Bench 4 measures emotional and social judgment in multi-turn chats. Creative Writing v3 assesses creative prose; Longform Writing assesses planning and consistency across eight chapters.
Does not tell you: Factual accuracy, every kind of editing, or performance with real users. These benchmarks use LLM judges and reflect their preferences. Scores come from the live leaderboard data, with distinct scales and coverage; individual run dates are not published.
Tells you: Whether people preferred the answer. Anonymous side-by-side votes turned into an Elo rating, in five arenas: general text, web development, vision, documents and search.
Does not tell you: Whether the answer was correct. Voters reward tone, length and formatting, and the text board applies style control to damp that but does not remove it. A gap smaller than the confidence interval is not a gap.
Tells you: Which models exist and what they cost. Context, input and output modalities, tool and schema support, per-provider price including cached reads, uptime, and the measured throughput and time-to-first-token behind each endpoint.
Does not tell you: Quality. It is a marketplace, so the catalog is what OpenRouter routes to — a model served only by its vendor’s own API may be missing. Throughput is a 30-minute window and moves.
Tells you: What the writing is actually like. Stock AI phrases and “not X, but Y” constructions per thousand words, repeated phrasing, vocabulary spread and reading grade — each counted the same way over 17 million words of human writing, so a model’s figure can be read against a person’s.
Does not tell you: Whether the writing is any good. It counts habits, not quality: a model can score well by being terse and dull. Shown on each model rather than folded into the ranking, because too few models have been measured for that.
Tells you: 500 real GitHub issues, graded by the maintainers’ own tests. The closest thing to “can it land a change in an existing codebase.”
Does not tell you: The current frontier. Entries are submitted by whoever built the scaffold, so the board rewards teams who bother to submit and runs months behind new releases.
Tells you: Whether a model left alone in a terminal finishes the job. A test suite decides, with no partial credit. The board also reports what each run cost and how long a trial took.
Does not tell you: The model on its own. A score belongs to a model and its scaffold — Codex, Claude Code — and swapping the harness moves the number. The board is small, so most of the catalog has no entry.
Sources considered but not used
For completeness, here are additional sources we considered using for benchmark data, but ultimately decided they were not quite right.
Useful document-understanding tasks, but the published table reviewed in September 2026 stops at July 2025 model releases. Coverage of current models is insufficient; preview and release variants cannot safely share scores.
Excellent, current coverage — but the data is commercially licensed. I do not want this page to depend on republishing or deriving from data without clear redistribution rights.