The models, ranked for people who have to ship with them
Twelve models that matter if you have to ship with them this quarter, on one scale. Independent evals carry half the weight, price a fifth, what practitioners say and how the thing behaves in production the rest. First edition, so everything is marked new; from here on, arrows show who moved. Every number links to who measured it and when Wren read it.
Cost against intelligence
Artificial Analysis Intelligence Index on the vertical, blended price per million tokens (three input to one output, log scale) on the horizontal. Dashed lines are the medians of this table. Top left is where you want to be.
Solid dots use vendor list prices; hollow dots use Artificial Analysis' blended estimate where the vendor does not publish a list price. Not plotted (no independent index yet): DeepSeek Flash (4.1).
The model people actually ship agents on: top of the independent index, a price that does not need a budget meeting, and a community that complains mainly about how much it has to be used.
The reasoning heavyweight: leads the hardest exams and the cipher-cracking party tricks, and charges like it.
Wins the popularity contest: first on Arena, mid-pack on the index, and priced like a workhorse.
The value pick of the table: second on the index at a fifth of the Opus price, and the fastest of the top five.
Near-Astra intelligence at a fifth of the price, which is OpenAI's own line and, for once, the index agrees.
The specialist: Anthropic's own docs say to reach for it when Opus 5.5 at high effort still falls short, and the price says the same.
Ten cents a million and, per the people using it, smarter than the other ten-cent model; mind the 100K cliff.
The fastest thing in the table by a distance, and the one to beat on the stress benchmark nobody else talks about.
The reason a thousand people argued on Hacker News this week: frontier-adjacent coding for pocket change, with the caveats you would expect.
OpenAI's ten-cent model: the one every ChatGPT Free user now has, and the one Haiku users keep comparing themselves to.
The highest-placed open-licence model on Arena, and the one most teams have not tried yet.
Mid-table on the index at a frontier price, with a smaller context window than everyone above it.
How the score works
Score is 0 to 100: 50% independent evals normalised within this table, 20% price-performance, 15% practitioner sentiment from Signals and community threads, 15% operational fit (context, speed, availability, open weights). Models within three points are a tie in practice. Vendor-published figures are labelled as such and never counted as independent.
Benchmarks are reported with who ran them and when Wren read them. Vendor figures are labelled as vendor figures. Sentiment comes from Signals and linked threads, attributed, not verified. Disagree with a rank? The open thread is the place.