{"data":{"updated_at":"2026-10-10T13:40:00.000Z","intro":"Twelve models that matter if you have to ship with them this quarter, on one scale. Independent evals carry half the weight, price a fifth, what practitioners say and how the thing behaves in production the rest. First edition, so everything is marked new; from here on, arrows show who moved. Every number links to who measured it and when Wren read it.","methodology":"Score is 0 to 100: 50% independent evals normalised within this table, 20% price-performance, 15% practitioner sentiment from Signals and community threads, 15% operational fit (context, speed, availability, open weights). Models within three points are a tie in practice. Vendor-published figures are labelled as such and never counted as independent.","models":[{"id":"claude-opus-5-5","name":"Claude Opus 5.5","vendor":"Anthropic","tier":"frontier","released":"2026-09-22","rank":1,"movement":"new","score":91,"verdict":"The model people actually ship agents on: top of the independent index, a price that does not need a budget meeting, and a community that complains mainly about how much it has to be used.","bestFor":["Agentic coding","Long-running agents","Enterprise knowledge work"],"caveats":["Arena puts Gemini 4 Argon ahead on chat preference; the gap is inside two scores' error bars."],"benchmarks":[{"name":"Artificial Analysis Intelligence Index (max)","value":"58","source":{"title":"Artificial Analysis LLM leaderboard","url":"https://artificialanalysis.ai/leaderboards/models","readAt":"2026-10-09T22:30:00Z"}},{"name":"Arena text score (high)","value":"1507","source":{"title":"Arena text leaderboard (Oct 8, 2026)","url":"https://arena.ai/leaderboard/chat/text","readAt":"2026-10-09T22:30:00Z"}},{"name":"Humanity's Last Exam, Diamond (Scale)","value":"55.0","source":{"title":"Scale SEAL leaderboards","url":"https://labs.scale.com/leaderboard","readAt":"2026-10-09T22:30:00Z"}},{"name":"Epoch Capabilities Index","value":"167","source":{"title":"Epoch AI benchmarks hub","url":"https://epoch.ai/benchmarks","readAt":"2026-10-09T22:30:00Z"}}],"price":{"input":4,"output":20,"currency":"USD","per":"1M tokens","source":{"title":"Claude pricing","url":"https://claude.com/pricing"}},"context":"1M","openWeights":false,"sentiment":{"summary":"The launch thread was the biggest Claude thread of the year on Hacker News, a port of the TypeScript compiler to Rust credited the model with a working build in ten hours, and a Vals AI materials-discovery run used it for agentic simulation. A 'has it been nerfed yet' tracker is the most-upvoted scepticism.","evidence":[{"title":"HN: Claude Opus 5.5 (1,806 points)","url":"https://news.ycombinator.com/item?id=49803892"},{"title":"HN: TypeScript compiler ported to Rust by an LLM","url":"https://news.ycombinator.com/item?id=50000676"},{"title":"HN: Livenerf, has Opus 5.5 been nerfed yet?","url":"https://news.ycombinator.com/item?id=49901736"}]},"articleSlug":"claude-haiku-5-5-opus-5-5-pricing"},{"id":"gpt-6-astra","name":"GPT-6 Astra","vendor":"OpenAI","tier":"frontier","released":"2026-09-03","rank":2,"movement":"new","score":86,"verdict":"The reasoning heavyweight: leads the hardest exams and the cipher-cracking party tricks, and charges like it.","bestFor":["Hard reasoning","Research-grade analysis","When cost is not the constraint"],"caveats":["Long-context pricing doubles above 272K input tokens.","Speed is the lowest of the frontier set in Artificial Analysis' measurement."],"benchmarks":[{"name":"Humanity's Last Exam, Diamond (Scale)","value":"60.6","source":{"title":"Scale SEAL leaderboards","url":"https://labs.scale.com/leaderboard","readAt":"2026-10-09T22:30:00Z"}},{"name":"Humanity's Sixth Sense (Scale)","value":"53.6","source":{"title":"Scale SEAL leaderboards","url":"https://labs.scale.com/leaderboard","readAt":"2026-10-09T22:30:00Z"}},{"name":"Artificial Analysis Intelligence Index (max)","value":"53","source":{"title":"Artificial Analysis LLM leaderboard","url":"https://artificialanalysis.ai/leaderboards/models","readAt":"2026-10-09T22:30:00Z"}},{"name":"FrontierMath Erdős (Epoch)","value":"3% (2 of 68)","source":{"title":"Epoch AI benchmarks hub","url":"https://epoch.ai/benchmarks","readAt":"2026-10-09T22:30:00Z"}}],"price":{"input":10,"output":50,"currency":"USD","per":"1M tokens","source":{"title":"OpenAI API pricing","url":"https://developers.openai.com/api/docs/pricing"}},"context":"1M","openWeights":false,"sentiment":{"summary":"The launch drew the largest model thread of 2026 on Hacker News, and the stories that followed were about capability rather than cost: an Enigma message unsolved since 2005, a WWI cipher, a driving benchmark.","evidence":[{"title":"HN: GPT-6 Astra (2,279 points)","url":"https://news.ycombinator.com/item?id=49554643"},{"title":"HN: GPT-6 Astra breaks an Enigma message unsolved since 2005","url":"https://news.ycombinator.com/item?id=49801324"}]}},{"id":"gemini-4-argon","name":"Gemini 4 Argon","vendor":"Google","tier":"frontier","released":"2026-09-30","rank":3,"movement":"new","score":85,"verdict":"Wins the popularity contest: first on Arena, mid-pack on the index, and priced like a workhorse.","bestFor":["Chat and assistant products","Google Cloud shops","Price-sensitive frontier work"],"caveats":["Google's public pricing page did not list Argon when read; the blended figure is Artificial Analysis' estimate."],"benchmarks":[{"name":"Arena text score (high)","value":"1525","source":{"title":"Arena text leaderboard (Oct 8, 2026)","url":"https://arena.ai/leaderboard/chat/text","readAt":"2026-10-09T22:30:00Z"}},{"name":"Artificial Analysis Intelligence Index (high)","value":"53","source":{"title":"Artificial Analysis LLM leaderboard","url":"https://artificialanalysis.ai/leaderboards/models","readAt":"2026-10-09T22:30:00Z"}},{"name":"Artificial Analysis blended price","value":"$1.99 per 1M","source":{"title":"Artificial Analysis LLM leaderboard","url":"https://artificialanalysis.ai/leaderboards/models","readAt":"2026-10-09T22:30:00Z"}}],"context":"1M","openWeights":false,"sentiment":{"summary":"The announcement thread ran to 1,190 comments; the independent analysis thread was smaller and more measured.","evidence":[{"title":"HN: Gemini 4 Argon (1,704 points)","url":"https://news.ycombinator.com/item?id=49913571"},{"title":"Artificial Analysis: Gemini 4 Argon analysis","url":"https://artificialanalysis.ai/models/gemini-4-argon"}]},"articleSlug":"google-one-agent-for-work-gemini-at-work-2026"},{"id":"claude-sonnet-5-5","name":"Claude Sonnet 5.5","vendor":"Anthropic","tier":"workhorse","rank":4,"movement":"new","score":84,"verdict":"The value pick of the table: second on the index at a fifth of the Opus price, and the fastest of the top five.","bestFor":["Production coding assistants","High-volume agents","Default model for most teams"],"benchmarks":[{"name":"Artificial Analysis Intelligence Index (max)","value":"56","source":{"title":"Artificial Analysis LLM leaderboard","url":"https://artificialanalysis.ai/leaderboards/models","readAt":"2026-10-09T22:30:00Z"}},{"name":"Artificial Analysis output speed (max)","value":"141 tokens/s","source":{"title":"Artificial Analysis LLM leaderboard","url":"https://artificialanalysis.ai/leaderboards/models","readAt":"2026-10-09T22:30:00Z"}}],"price":{"input":2,"output":10,"currency":"USD","per":"1M tokens","source":{"title":"Claude pricing","url":"https://claude.com/pricing"}},"context":"1M","openWeights":false,"sentiment":{"summary":"Less discussed than its siblings because it just works; the October cache-read price cut to $0.10 per million was the news.","evidence":[{"title":"HN: Claude Haiku 5.5 thread (Sonnet pricing discussed)","url":"https://news.ycombinator.com/item?id=49996437"}]},"articleSlug":"claude-haiku-5-5-opus-5-5-pricing"},{"id":"gpt-6-1-sol","name":"GPT-6.1 Sol","vendor":"OpenAI","tier":"workhorse","released":"2026-09-29","rank":5,"movement":"new","score":82,"verdict":"Near-Astra intelligence at a fifth of the price, which is OpenAI's own line and, for once, the index agrees.","bestFor":["OpenAI-standardised teams","Agents API and Decisions API","Cost-controlled reasoning"],"benchmarks":[{"name":"Artificial Analysis Intelligence Index (max)","value":"52","source":{"title":"Artificial Analysis LLM leaderboard","url":"https://artificialanalysis.ai/leaderboards/models","readAt":"2026-10-09T22:30:00Z"}},{"name":"Artificial Analysis blended price (max)","value":"$0.72 per 1M","source":{"title":"Artificial Analysis LLM leaderboard","url":"https://artificialanalysis.ai/leaderboards/models","readAt":"2026-10-09T22:30:00Z"}},{"name":"Humanity's Sixth Sense (Scale)","value":"46.6","source":{"title":"Scale SEAL leaderboards","url":"https://labs.scale.com/leaderboard","readAt":"2026-10-09T22:30:00Z"}}],"price":{"input":2,"output":10,"currency":"USD","per":"1M tokens","source":{"title":"OpenAI API pricing","url":"https://developers.openai.com/api/docs/pricing"}},"context":"1M","openWeights":false,"sentiment":{"summary":"The launch thread's title did the arguing: 'near-Astra for a fifth of the price' drew 955 comments, most of them about routing.","evidence":[{"title":"HN: GPT 6.1 Sol (1,067 points)","url":"https://news.ycombinator.com/item?id=49896586"}]},"articleSlug":"openai-devday-2026-agents-api-computer-use"},{"id":"claude-fable-5-1","name":"Claude Fable 5.1","vendor":"Anthropic","tier":"frontier","rank":6,"movement":"new","score":80,"verdict":"The specialist: Anthropic's own docs say to reach for it when Opus 5.5 at high effort still falls short, and the price says the same.","bestFor":["Long-horizon agents","Problems Opus fails on","Demanding reasoning"],"caveats":["Ten dollars per million input is the highest in this table alongside Astra."],"benchmarks":[{"name":"Artificial Analysis Intelligence Index (max)","value":"53","source":{"title":"Artificial Analysis LLM leaderboard","url":"https://artificialanalysis.ai/leaderboards/models","readAt":"2026-10-09T22:30:00Z"}},{"name":"Arena text score (max)","value":"1501","source":{"title":"Arena text leaderboard (Oct 8, 2026)","url":"https://arena.ai/leaderboard/chat/text","readAt":"2026-10-09T22:30:00Z"}},{"name":"Humanity's Last Exam, Diamond (Scale)","value":"51.3","source":{"title":"Scale SEAL leaderboards","url":"https://labs.scale.com/leaderboard","readAt":"2026-10-09T22:30:00Z"}}],"price":{"input":10,"output":50,"currency":"USD","per":"1M tokens","source":{"title":"Claude pricing","url":"https://claude.com/pricing"}},"context":"1M","openWeights":false},{"id":"claude-haiku-5-5","name":"Claude Haiku 5.5","vendor":"Anthropic","tier":"fast","released":"2026-10-07","rank":7,"movement":"new","score":74,"verdict":"Ten cents a million and, per the people using it, smarter than the other ten-cent model; mind the 100K cliff.","bestFor":["Classification and extraction","Routing tiers","Latency-sensitive paths"],"caveats":["Price jumps five-fold above 100K input tokens, which commenters call an absurdly low cutoff for agent runs."],"benchmarks":[{"name":"Vendor list price under 100K tokens","value":"$0.10 in / $0.50 out","source":{"title":"Claude pricing","url":"https://claude.com/pricing","readAt":"2026-10-09T22:30:00Z"}},{"name":"Artificial Analysis Intelligence Index (max)","value":"43","source":{"title":"Artificial Analysis LLM leaderboard","url":"https://artificialanalysis.ai/leaderboards/models","readAt":"2026-10-10T13:30:00Z"}},{"name":"Artificial Analysis output speed (max)","value":"240 tokens/s","source":{"title":"Artificial Analysis LLM leaderboard","url":"https://artificialanalysis.ai/leaderboards/models","readAt":"2026-10-10T13:30:00Z"}}],"price":{"input":0.1,"output":0.5,"currency":"USD","per":"1M tokens","source":{"title":"Claude pricing","url":"https://claude.com/pricing"}},"context":"1M","openWeights":false,"sentiment":{"summary":"The thread had over a thousand points in a day; the recurring reports are near-perfect structured classification in production and 'noticeably smarter than GPT-6 Luna', against complaints about the 100K pricing tier.","evidence":[{"title":"HN: Claude Haiku 5.5 (1,043 points)","url":"https://news.ycombinator.com/item?id=49996437"}]},"articleSlug":"claude-haiku-5-5-opus-5-5-pricing"},{"id":"muse-spark-1-3","name":"Muse Spark 1.3","vendor":"Meta","tier":"workhorse","rank":8,"movement":"new","score":72,"verdict":"The fastest thing in the table by a distance, and the one to beat on the stress benchmark nobody else talks about.","bestFor":["Latency-first products","Customer-facing chat","Meta-ecosystem teams"],"benchmarks":[{"name":"Artificial Analysis output speed (max)","value":"175 tokens/s","source":{"title":"Artificial Analysis LLM leaderboard","url":"https://artificialanalysis.ai/leaderboards/models","readAt":"2026-10-09T22:30:00Z"}},{"name":"Artificial Analysis Intelligence Index (max)","value":"48","source":{"title":"Artificial Analysis LLM leaderboard","url":"https://artificialanalysis.ai/leaderboards/models","readAt":"2026-10-09T22:30:00Z"}},{"name":"Arena text score (max)","value":"1494","source":{"title":"Arena text leaderboard (Oct 8, 2026)","url":"https://arena.ai/leaderboard/chat/text","readAt":"2026-10-09T22:30:00Z"}},{"name":"DistressBench (Scale)","value":"84.88","source":{"title":"Scale SEAL leaderboards","url":"https://labs.scale.com/leaderboard","readAt":"2026-10-09T22:30:00Z"}},{"name":"Artificial Analysis blended price (max)","value":"$0.78 per 1M","source":{"title":"Artificial Analysis LLM leaderboard","url":"https://artificialanalysis.ai/leaderboards/models","readAt":"2026-10-10T13:30:00Z"}}]},{"id":"deepseek-flash","name":"DeepSeek Flash (4.1)","vendor":"DeepSeek","tier":"fast","rank":9,"movement":"new","score":70,"verdict":"The reason a thousand people argued on Hacker News this week: frontier-adjacent coding for pocket change, with the caveats you would expect.","bestFor":["Cheap coding tiers","Bulk generation","Experiments where data residency is not a constraint"],"caveats":["Open-weights status and licence not verified from a primary source at time of writing.","Off-peak and peak pricing differ; the cheap numbers are off-peak.","Data handling terms should be read before any enterprise use."],"benchmarks":[{"name":"Vendor list price, cache miss (off-peak to peak)","value":"$0.15 to $0.30 in / $0.60 to $1.20 out","source":{"title":"DeepSeek API pricing","url":"https://api-docs.deepseek.com/quick_start/pricing","readAt":"2026-10-09T22:30:00Z"}}],"price":{"input":0.3,"output":1.2,"currency":"USD","per":"1M tokens","source":{"title":"DeepSeek API pricing","url":"https://api-docs.deepseek.com/quick_start/pricing"}},"context":"1M","sentiment":{"summary":"The 'why isn't the industry freaking out' post hit 1,067 points; the most-upvoted replies say the price is real and the coding gap to Opus is also real.","evidence":[{"title":"HN: Why isn't the industry freaking out about DeepSeek 4.1 Flash?","url":"https://news.ycombinator.com/item?id=50000488"}]}},{"id":"gpt-6-luna","name":"GPT-6 Luna","vendor":"OpenAI","tier":"fast","rank":10,"movement":"new","score":66,"verdict":"OpenAI's ten-cent model: the one every ChatGPT Free user now has, and the one Haiku users keep comparing themselves to.","bestFor":["ChatGPT Free and Go tier","Simple extraction","Cost floors"],"benchmarks":[{"name":"Vendor list price","value":"$0.10 in / $0.50 out","source":{"title":"OpenAI API pricing","url":"https://developers.openai.com/api/docs/pricing","readAt":"2026-10-09T22:30:00Z"}},{"name":"Artificial Analysis Intelligence Index","value":"38","source":{"title":"Artificial Analysis LLM leaderboard","url":"https://artificialanalysis.ai/leaderboards/models","readAt":"2026-10-10T13:30:00Z"}},{"name":"Artificial Analysis output speed","value":"137 tokens/s","source":{"title":"Artificial Analysis LLM leaderboard","url":"https://artificialanalysis.ai/leaderboards/models","readAt":"2026-10-10T13:30:00Z"}}],"price":{"input":0.1,"output":0.5,"currency":"USD","per":"1M tokens","source":{"title":"OpenAI API pricing","url":"https://developers.openai.com/api/docs/pricing"}},"context":"1M","openWeights":false,"articleSlug":"gpt-6-chatgpt-intelligent-ui"},{"id":"kimi-k3","name":"Kimi K3","vendor":"Moonshot","tier":"open","rank":11,"movement":"new","score":60,"verdict":"The highest-placed open-licence model on Arena, and the one most teams have not tried yet.","bestFor":["Self-hosting","Open-licence requirements","Arena-style chat quality"],"caveats":["Price is Artificial Analysis' blended figure, not a vendor list price; context not verified from a primary source."],"benchmarks":[{"name":"Arena text score (max)","value":"1488","source":{"title":"Arena text leaderboard (Oct 8, 2026)","url":"https://arena.ai/leaderboard/chat/text","readAt":"2026-10-09T22:30:00Z"}},{"name":"Artificial Analysis Intelligence Index (max)","value":"44","source":{"title":"Artificial Analysis LLM leaderboard","url":"https://artificialanalysis.ai/leaderboards/models","readAt":"2026-10-10T13:30:00Z"}},{"name":"Artificial Analysis blended price (max)","value":"$2.31 per 1M","source":{"title":"Artificial Analysis LLM leaderboard","url":"https://artificialanalysis.ai/leaderboards/models","readAt":"2026-10-10T13:30:00Z"}}],"openWeights":true},{"id":"grok-4-7","name":"Grok 4.7","vendor":"xAI","tier":"workhorse","rank":12,"movement":"new","score":58,"verdict":"Mid-table on the index at a frontier price, with a smaller context window than everyone above it.","bestFor":["X-integrated products"],"caveats":["500K context is the smallest in the table.","Listed by Artificial Analysis under the creator's current corporate name."],"benchmarks":[{"name":"Artificial Analysis Intelligence Index (xhigh)","value":"46","source":{"title":"Artificial Analysis LLM leaderboard","url":"https://artificialanalysis.ai/leaderboards/models","readAt":"2026-10-09T22:30:00Z"}},{"name":"Artificial Analysis blended price (xhigh)","value":"$3.74 per 1M","source":{"title":"Artificial Analysis LLM leaderboard","url":"https://artificialanalysis.ai/leaderboards/models","readAt":"2026-10-09T22:30:00Z"}}],"context":"500K","openWeights":false}],"content_hash":"77f76ca8e61bfe88bf40e2473a2cda2734d67984e74d1e72ac31ff920c84b63d"},"meta":{"site":"Models at Work","canonical":"https://modelsatwork.news","generated_at":"2026-10-11T00:50:02.102Z","license":"CC BY 4.0","license_url":"https://creativecommons.org/licenses/by/4.0/","agents_guide":"https://modelsatwork.news/agents"}}