GPT-6 Luna vs Mistral Small 4
GPT-6 Luna scores higher on the Models at Work leaderboard (74 vs 42 out of 100). GPT-6 Luna is about 1.3× cheaper per blended million tokens ($0.20 vs $0.26). On Artificial Analysis' independent Intelligence Index, GPT-6 Luna leads 38 to 11.
| GPT-6 Luna · OpenAI | Mistral Small 4 · Mistral AI | |
|---|---|---|
| Leaderboard rank | #13 | #24 |
| Score (0–100) | 74 | 42 |
| Tier | Fast and cheap | Open weights |
| Price per 1M tokens | $0.10 in · $0.50 out | $0.15 in · $0.60 out |
| Context window | 1M | 256K |
| Open weights | No | Yes |
| Artificial Analysis Intelligence Index | 38 | 11 |
| Best for | ChatGPT Free and Go tier, Simple extraction, Cost floors | Self-hosted multimodal chat on one GPU, Fine-tuning base for EU-regulated teams, Cheap fast tasks with a reasoning_effort dial |
Choose GPT-6 Luna if…
- you need chatgpt free and go tier
- you need simple extraction
- you need cost floors
OpenAI's ten-cent model: the one every ChatGPT Free user now has, and the one Haiku users keep comparing themselves to.
Choose Mistral Small 4 if…
- you need self-hosted multimodal chat on one gpu
- you need fine-tuning base for eu-regulated teams
- you need cheap fast tasks with a reasoning_effort dial
One Apache-2.0 model replacing Magistral, Pixtral and Devstral: 119B MoE with 6B active, so it fits one H100 at 4-bit and runs at ~165 tok/s. $0.15/$0.60 is hard to beat for a self-hostable multimodal model, but it is not clever.
Caveats: AA 11; HN testers rate it 'okay, nothing exceptional' and 'not the best' for complex agentic work. Not listed on Arena; limited independent evaluation. Mistral Medium 3.5 (Modified MIT) is the step up if you need more intelligence from the same vendor.
What practitioners say
Mistral Small 4: Modest HN reception (126 points). Fans liked the economics: '$0.60/1M output is a steal' versus Qwen models that 'waste tokens on reasoning', and that ~120B 'fits onto a single H100 with 4 bit quant'. Others were blunt: 'okay, nothing exceptional', 'I'd use it for some basic tasks but not actual complex tasks', and one tester called it 'worse than glm air 4.5'. A later comment found the 4-bit quant roughly comparable to Qwen 3.6 27B.
Questions people ask
Which is better, GPT-6 Luna or Mistral Small 4?
GPT-6 Luna scores higher on the Models at Work leaderboard (74 vs 42 out of 100). GPT-6 Luna is about 1.3× cheaper per blended million tokens ($0.20 vs $0.26). On Artificial Analysis' independent Intelligence Index, GPT-6 Luna leads 38 to 11.
When should I choose GPT-6 Luna over Mistral Small 4?
Choose GPT-6 Luna for chatgpt free and go tier, simple extraction, cost floors. OpenAI's ten-cent model: the one every ChatGPT Free user now has, and the one Haiku users keep comparing themselves to.
When should I choose Mistral Small 4 over GPT-6 Luna?
Choose Mistral Small 4 for self-hosted multimodal chat on one gpu, fine-tuning base for eu-regulated teams, cheap fast tasks with a reasoning_effort dial. One Apache-2.0 model replacing Magistral, Pixtral and Devstral: 119B MoE with 6B active, so it fits one H100 at 4-bit and runs at ~165 tok/s. $0.15/$0.60 is hard to beat for a self-hostable multimodal model, but it is not clever.
Which is cheaper, GPT-6 Luna or Mistral Small 4?
GPT-6 Luna costs $0.10 per million input tokens and $0.50 per million output tokens on OpenAI's list price (OpenAI API pricing, read Oct 11, 2026). Mistral Small 4 costs $0.15 per million input tokens and $0.60 per million output tokens on Mistral AI's list price (Mistral Small 4 announcement, read Oct 11, 2026).
Scores come from the Models at Work leaderboard: Score is 0 to 100 and computed, not typed: 55% independent evals (Artificial Analysis, Arena, Scale SEAL, Epoch, each scaled against its natural floor and the best score in this table, then averaged), 15% blended price on a fixed log scale ($0.05 per million is 100, $60 is 0), 15% practitioner sentiment from Signals and community threads, 15% operational fit (context window, open weights). A missing component drops out and the rest are reweighted. A model with no independent eval yet is provisional and ranks below every measured one. Within three points is a tie. Vendor-published figures are labelled and never counted. Prices are vendor list prices where published; blended figures assume three input tokens per output token.