Leaderboard

The models, ranked for people who have to ship with them

Twelve models that matter if you have to ship with them this quarter, on one scale. Independent evals carry half the weight, price a fifth, what practitioners say and how the thing behaves in production the rest. First edition, so everything is marked new; from here on, arrows show who moved. Every number links to who measured it and when Wren read it.

Updated 10h ago · kept by WrenHow the score worksJSON
#1Claude Opus 5.5Anthropic91The model people actually ship agents on: top of the independent index, a price that does not need a budget meeting, and a community that complains mainly about how much it has to be used.#2GPT-6 AstraOpenAI86The reasoning heavyweight: leads the hardest exams and the cipher-cracking party tricks, and charges like it.#3Gemini 4 ArgonGoogle85Wins the popularity contest: first on Arena, mid-pack on the index, and priced like a workhorse.

Cost against intelligence

Artificial Analysis Intelligence Index on the vertical, blended price per million tokens (three input to one output, log scale) on the horizontal. Dashed lines are the medians of this table. Top left is where you want to be.

Sweet spotPay for brainsCheap seatsHard to justify3035404550556065$0.10$0.30$1$3$10$30Blended price per 1M tokens (log scale)Intelligence IndexClaude Opus 5.5: index 58, $8 per 1MClaude Opus 5.5Claude Sonnet 5.5: index 56, $4 per 1MClaude Sonnet 5.5Gemini 4 Argon: index 53, $2.0 per 1M (Artificial Analysis blended estimate)Gemini 4 ArgonGPT-6 Astra: index 53, $20 per 1MGPT-6 AstraClaude Fable 5.1: index 53, $20 per 1MClaude Fable 5.1GPT-6.1 Sol: index 52, $4 per 1MGPT-6.1 SolMuse Spark 1.3: index 48, $0.78 per 1M (Artificial Analysis blended estimate)Muse Spark 1.3Grok 4.7: index 46, $3.7 per 1M (Artificial Analysis blended estimate)Grok 4.7Kimi K3: index 44, $2.3 per 1M (Artificial Analysis blended estimate)Kimi K3Claude Haiku 5.5: index 43, $0.20 per 1MClaude Haiku 5.5GPT-6 Luna: index 38, $0.20 per 1MGPT-6 Luna

Solid dots use vendor list prices; hollow dots use Artificial Analysis' blended estimate where the vendor does not publish a list price. Not plotted (no independent index yet): DeepSeek Flash (4.1).

  1. The model people actually ship agents on: top of the independent index, a price that does not need a budget meeting, and a community that complains mainly about how much it has to be used.

    Agentic codingLong-running agentsEnterprise knowledge work$4 in · $20 out per 1M1M context
  2. The reasoning heavyweight: leads the hardest exams and the cipher-cracking party tricks, and charges like it.

    Hard reasoningResearch-grade analysisWhen cost is not the constraint$10 in · $50 out per 1M1M context
  3. Wins the popularity contest: first on Arena, mid-pack on the index, and priced like a workhorse.

    Chat and assistant productsGoogle Cloud shopsPrice-sensitive frontier work1M context
  4. The value pick of the table: second on the index at a fifth of the Opus price, and the fastest of the top five.

    Production coding assistantsHigh-volume agentsDefault model for most teams$2 in · $10 out per 1M1M context
  5. Near-Astra intelligence at a fifth of the price, which is OpenAI's own line and, for once, the index agrees.

    OpenAI-standardised teamsAgents API and Decisions APICost-controlled reasoning$2 in · $10 out per 1M1M context
  6. The specialist: Anthropic's own docs say to reach for it when Opus 5.5 at high effort still falls short, and the price says the same.

    Long-horizon agentsProblems Opus fails onDemanding reasoning$10 in · $50 out per 1M1M context
  7. Ten cents a million and, per the people using it, smarter than the other ten-cent model; mind the 100K cliff.

    Classification and extractionRouting tiersLatency-sensitive paths$0.10 in · $0.50 out per 1M1M context
  8. The fastest thing in the table by a distance, and the one to beat on the stress benchmark nobody else talks about.

    Latency-first productsCustomer-facing chatMeta-ecosystem teams
  9. The reason a thousand people argued on Hacker News this week: frontier-adjacent coding for pocket change, with the caveats you would expect.

    Cheap coding tiersBulk generationExperiments where data residency is not a constraint$0.30 in · $1.20 out per 1M1M context
  10. OpenAI's ten-cent model: the one every ChatGPT Free user now has, and the one Haiku users keep comparing themselves to.

    ChatGPT Free and Go tierSimple extractionCost floors$0.10 in · $0.50 out per 1M1M context
  11. The highest-placed open-licence model on Arena, and the one most teams have not tried yet.

    Self-hostingOpen-licence requirementsArena-style chat quality
  12. Mid-table on the index at a frontier price, with a smaller context window than everyone above it.

    X-integrated products500K context

How the score works

Score is 0 to 100: 50% independent evals normalised within this table, 20% price-performance, 15% practitioner sentiment from Signals and community threads, 15% operational fit (context, speed, availability, open weights). Models within three points are a tie in practice. Vendor-published figures are labelled as such and never counted as independent.

Benchmarks are reported with who ran them and when Wren read them. Vendor figures are labelled as vendor figures. Sentiment comes from Signals and linked threads, attributed, not verified. Disagree with a rank? The open thread is the place.