Leaderboard · #22 of 26 · Google · Open weights

Gemma 4 31B

The dense Apache-2.0 model local-first teams call 'the new baseline'. Excellent at rule-following, OCR and automation; needs ~48GB for full 256K context and is not a frontier reasoner. Free to run, no API price to speak of.

Updated Oct 11, 2026 · kept by WrenJSON

Key facts

Score
58 / 100, rank #22
Evals47Pricen/aPeople72Fit85
Price
No verified list price
Context window
256K
Open weights
Yes
Released
Apr 2, 2026
Best for
On-device and workstation inference, OCR and document vision, Deterministic pipeline/automation steps

Google has not published a list price for Gemma 4 31B that we could verify from a primary source.

Independent evals

BenchmarkResultWho ran itRead
Artificial Analysis Intelligence Index15Gemma 4 31B - Artificial AnalysisOct 10, 2026
Arena Text score1452 (rank 75)Text Arena LeaderboardOct 10, 2026
  • Artificial Analysis output speed (tokens/s): 34.8 (Gemma 4 31B - Artificial Analysis)

What practitioners say

HN's local-LLM crowd treats it as the reference point: 'the new baseline for local models' (soganess, 70GB peak on an M5 Max at 256K context); 'particularly good at pipeline/automation tasks' and better than Qwen even at 100B+ for rule-following; 'very good at OCR'. One user had it catch a bug Opus 4.7 missed. Counterpoints: Qwen 3.6 'handles context a little better', peers 'predominantly run Qwen', and ~11-16 tok/s on Macs feels slow.

Caveats

  • Dense 31B: expect ~12-16 tok/s on Apple Silicon at Q4-Q6, and 70GB RAM at full context.
  • Several HN users say Qwen 3.6/3.8 handles long context and agentic tasks a little better; Gemma wins on instruction discipline.
  • AA 15 on the current index; this is a local model, not a hosted-API competitor.

Questions people ask

How much does Gemma 4 31B cost?

Google has not published a list price for Gemma 4 31B that we could verify from a primary source.

How good is Gemma 4 31B?

Gemma 4 31B ranks #22 of 26 on the Models at Work leaderboard with a score of 58 out of 100 (October 2026 edition). The dense Apache-2.0 model local-first teams call 'the new baseline'. Excellent at rule-following, OCR and automation; needs ~48GB for full 256K context and is not a frontier reasoner. Free to run, no API price to speak of.

What is Gemma 4 31B best for?

Gemma 4 31B is best for on-device and workstation inference, ocr and document vision, deterministic pipeline/automation steps.

What are Gemma 4 31B's benchmark scores?

Artificial Analysis Intelligence Index: 15 (Gemma 4 31B - Artificial Analysis); Arena Text score: 1452 (rank 75) (Text Arena Leaderboard).

What is Gemma 4 31B's context window?

Gemma 4 31B has a 256K token context window according to Google.

Is Gemma 4 31B open weights?

Yes. Gemma 4 31B is released with open weights, so it can be self-hosted.

Compare Gemma 4 31B

Methodology: Score is 0 to 100 and computed, not typed: 55% independent evals (Artificial Analysis, Arena, Scale SEAL, Epoch, each scaled against its natural floor and the best score in this table, then averaged), 15% blended price on a fixed log scale ($0.05 per million is 100, $60 is 0), 15% practitioner sentiment from Signals and community threads, 15% operational fit (context window, open weights). A missing component drops out and the rest are reweighted. A model with no independent eval yet is provisional and ranks below every measured one. Within three points is a tie. Vendor-published figures are labelled and never counted.