Head to head · October 2026

Mistral Large 4 vs Mistral Small 4

Mistral Large 4 scores higher on the Models at Work leaderboard (61 vs 42 out of 100). Mistral Small 4 is about 7.9× cheaper per blended million tokens ($0.26 vs $2.06). On Artificial Analysis' independent Intelligence Index, Mistral Large 4 leads 38 to 11.

Updated Oct 11, 2026 · every number links to who measured it
Mistral Large 4 · Mistral AIMistral Small 4 · Mistral AI
Leaderboard rank#20#24
Score (0–100)6142
TierWorkhorseOpen weights
Price per 1M tokens$1.36 in · $4.18 out$0.15 in · $0.60 out
Context window524K256K
Open weightsNoYes
Artificial Analysis Intelligence Index3811
Arena Text score1429 (rank 115)—
DeepSWE v1.1 (vendor)61.7%—
Best forEU-hosted or sovereignty-constrained deployments, Cybersecurity tooling, Vision-grounded document tasksSelf-hosted multimodal chat on one GPU, Fine-tuning base for EU-regulated teams, Cheap fast tasks with a reasoning_effort dial

Choose Mistral Large 4 if…

  • you need eu-hosted or sovereignty-constrained deployments
  • you need cybersecurity tooling
  • you need vision-grounded document tasks

Europe's 1T-parameter answer, four days old and still a preview. Independent numbers put it mid-pack (AA 38, Arena 1429) despite strong vendor coding and cyber claims. Pick it for EU hosting and policy reasons, not raw capability.

Caveats: Weights are promised for 27 Oct under an as-yet-unnamed custom licence; today it is API-only. Large 3 was Apache 2.0, so do not assume the same. Preview pricing is confusing: list is $1.36/$4.18 but the model page was showing $0.68/$2.09 at launch with no stated end date. Only two reasoning settings (none/high) and HN found little difference between them.

Choose Mistral Small 4 if…

  • you need self-hosted multimodal chat on one gpu
  • you need fine-tuning base for eu-regulated teams
  • you need cheap fast tasks with a reasoning_effort dial

One Apache-2.0 model replacing Magistral, Pixtral and Devstral: 119B MoE with 6B active, so it fits one H100 at 4-bit and runs at ~165 tok/s. $0.15/$0.60 is hard to beat for a self-hostable multimodal model, but it is not clever.

Caveats: AA 11; HN testers rate it 'okay, nothing exceptional' and 'not the best' for complex agentic work. Not listed on Arena; limited independent evaluation. Mistral Medium 3.5 (Modified MIT) is the step up if you need more intelligence from the same vendor.

What practitioners say

Mistral Large 4: Enormous HN reaction (2,037 points, 1,211 comments) split down the middle. Fans like the European training run (3,800 Grace Blackwell GPUs in Mistral's own datacentres), 'basically instant responses', and that it beats Chinese models on cyber benchmarks. Critics called it 'mediocre' and months late, and oh_no noted Mistral's published GDP.pdf numbers do not line up with Artificial Analysis's independent run. simonw found reasoning 'high' barely changes output.

Mistral Small 4: Modest HN reception (126 points). Fans liked the economics: '$0.60/1M output is a steal' versus Qwen models that 'waste tokens on reasoning', and that ~120B 'fits onto a single H100 with 4 bit quant'. Others were blunt: 'okay, nothing exceptional', 'I'd use it for some basic tasks but not actual complex tasks', and one tester called it 'worse than glm air 4.5'. A later comment found the 4-bit quant roughly comparable to Qwen 3.6 27B.

Questions people ask

Which is better, Mistral Large 4 or Mistral Small 4?

Mistral Large 4 scores higher on the Models at Work leaderboard (61 vs 42 out of 100). Mistral Small 4 is about 7.9× cheaper per blended million tokens ($0.26 vs $2.06). On Artificial Analysis' independent Intelligence Index, Mistral Large 4 leads 38 to 11.

When should I choose Mistral Large 4 over Mistral Small 4?

Choose Mistral Large 4 for eu-hosted or sovereignty-constrained deployments, cybersecurity tooling, vision-grounded document tasks. Europe's 1T-parameter answer, four days old and still a preview. Independent numbers put it mid-pack (AA 38, Arena 1429) despite strong vendor coding and cyber claims. Pick it for EU hosting and policy reasons, not raw capability.

When should I choose Mistral Small 4 over Mistral Large 4?

Choose Mistral Small 4 for self-hosted multimodal chat on one gpu, fine-tuning base for eu-regulated teams, cheap fast tasks with a reasoning_effort dial. One Apache-2.0 model replacing Magistral, Pixtral and Devstral: 119B MoE with 6B active, so it fits one H100 at 4-bit and runs at ~165 tok/s. $0.15/$0.60 is hard to beat for a self-hostable multimodal model, but it is not clever.

Which is cheaper, Mistral Large 4 or Mistral Small 4?

Mistral Large 4 costs $1.36 per million input tokens and $4.18 per million output tokens on Mistral AI's list price (Mistral Large 4 announcement, read Oct 11, 2026). Mistral Small 4 costs $0.15 per million input tokens and $0.60 per million output tokens on Mistral AI's list price (Mistral Small 4 announcement, read Oct 11, 2026).

Scores come from the Models at Work leaderboard: Score is 0 to 100 and computed, not typed: 55% independent evals (Artificial Analysis, Arena, Scale SEAL, Epoch, each scaled against its natural floor and the best score in this table, then averaged), 15% blended price on a fixed log scale ($0.05 per million is 100, $60 is 0), 15% practitioner sentiment from Signals and community threads, 15% operational fit (context window, open weights). A missing component drops out and the rest are reweighted. A model with no independent eval yet is provisional and ranks below every measured one. Within three points is a tie. Vendor-published figures are labelled and never counted. Prices are vendor list prices where published; blended figures assume three input tokens per output token.