Head to head · October 2026

Gemini 4 Argon vs Gemma 4 31B

Gemini 4 Argon scores higher on the Models at Work leaderboard (85 vs 58 out of 100). On Artificial Analysis' independent Intelligence Index, Gemini 4 Argon leads 53 to 15.

Updated Oct 11, 2026 · every number links to who measured it
Gemini 4 Argon · GoogleGemma 4 31B · Google
Leaderboard rank#3#22
Score (0–100)8558
TierFrontierOpen weights
Price per 1M tokens$1.99 blended (estimate)Not verified
Context window1M256K
Open weightsNoYes
Artificial Analysis Intelligence Index5315
Arena text score1525—
Arena Text score—1452 (rank 75)
Best forChat and assistant products, Google Cloud shops, Price-sensitive frontier workOn-device and workstation inference, OCR and document vision, Deterministic pipeline/automation steps

Choose Gemini 4 Argon if…

  • you need chat and assistant products
  • you need google cloud shops
  • you need price-sensitive frontier work

Wins the popularity contest: first on Arena, mid-pack on the index, and priced like a workhorse.

Caveats: Google's public pricing page did not list Argon when read; the blended figure is Artificial Analysis' estimate.

Choose Gemma 4 31B if…

  • you need on-device and workstation inference
  • you need ocr and document vision
  • you need deterministic pipeline/automation steps

The dense Apache-2.0 model local-first teams call 'the new baseline'. Excellent at rule-following, OCR and automation; needs ~48GB for full 256K context and is not a frontier reasoner. Free to run, no API price to speak of.

Caveats: Dense 31B: expect ~12-16 tok/s on Apple Silicon at Q4-Q6, and 70GB RAM at full context. Several HN users say Qwen 3.6/3.8 handles long context and agentic tasks a little better; Gemma wins on instruction discipline. AA 15 on the current index; this is a local model, not a hosted-API competitor.

What practitioners say

Gemini 4 Argon: The announcement thread ran to 1,190 comments; the independent analysis thread was smaller and more measured.

Gemma 4 31B: HN's local-LLM crowd treats it as the reference point: 'the new baseline for local models' (soganess, 70GB peak on an M5 Max at 256K context); 'particularly good at pipeline/automation tasks' and better than Qwen even at 100B+ for rule-following; 'very good at OCR'. One user had it catch a bug Opus 4.7 missed. Counterpoints: Qwen 3.6 'handles context a little better', peers 'predominantly run Qwen', and ~11-16 tok/s on Macs feels slow.

Questions people ask

Which is better, Gemini 4 Argon or Gemma 4 31B?

Gemini 4 Argon scores higher on the Models at Work leaderboard (85 vs 58 out of 100). On Artificial Analysis' independent Intelligence Index, Gemini 4 Argon leads 53 to 15.

When should I choose Gemini 4 Argon over Gemma 4 31B?

Choose Gemini 4 Argon for chat and assistant products, google cloud shops, price-sensitive frontier work. Wins the popularity contest: first on Arena, mid-pack on the index, and priced like a workhorse.

When should I choose Gemma 4 31B over Gemini 4 Argon?

Choose Gemma 4 31B for on-device and workstation inference, ocr and document vision, deterministic pipeline/automation steps. The dense Apache-2.0 model local-first teams call 'the new baseline'. Excellent at rule-following, OCR and automation; needs ~48GB for full 256K context and is not a frontier reasoner. Free to run, no API price to speak of.

Scores come from the Models at Work leaderboard: Score is 0 to 100 and computed, not typed: 55% independent evals (Artificial Analysis, Arena, Scale SEAL, Epoch, each scaled against its natural floor and the best score in this table, then averaged), 15% blended price on a fixed log scale ($0.05 per million is 100, $60 is 0), 15% practitioner sentiment from Signals and community threads, 15% operational fit (context window, open weights). A missing component drops out and the rest are reweighted. A model with no independent eval yet is provisional and ranks below every measured one. Within three points is a tie. Vendor-published figures are labelled and never counted. Prices are vendor list prices where published; blended figures assume three input tokens per output token.