Gemini 3.5 Flash-Lite vs Gemma 4 31B
Gemini 3.5 Flash-Lite scores higher on the Models at Work leaderboard (63 vs 58 out of 100). On Artificial Analysis' independent Intelligence Index, Gemini 3.5 Flash-Lite leads 22 to 15.
| Gemini 3.5 Flash-Lite · Google | Gemma 4 31B · Google | |
|---|---|---|
| Leaderboard rank | #19 | #22 |
| Score (0–100) | 63 | 58 |
| Tier | Fast and cheap | Open weights |
| Price per 1M tokens | $0.30 in · $2.50 out | Not verified |
| Context window | 1M | 256K |
| Open weights | No | Yes |
| Artificial Analysis Intelligence Index | 22 | 15 |
| Arena Text score | 1455 (rank 70) | 1452 (rank 75) |
| Best for | Classification and extraction pipelines, Routing and triage, Agentic search at volume | On-device and workstation inference, OCR and document vision, Deterministic pipeline/automation steps |
Choose Gemini 3.5 Flash-Lite if…
- you need classification and extraction pipelines
- you need routing and triage
- you need agentic search at volume
Google's volume tier: 1M context, ~430 tok/s, $0.30/$2.50, and a respectable Arena rank 70. The right default for classification, extraction and routing at scale. Not a reasoning model in any serious sense (AA 22).
Caveats: Output at $2.50/1M is pricier than GPT-6 Luna or GLM-5.3-Flash for the intelligence you get; the win is speed, not cost per token. Google has shipped three Flash-Lite generations in 2026; plan for deprecation cycles.
Choose Gemma 4 31B if…
- you need on-device and workstation inference
- you need ocr and document vision
- you need deterministic pipeline/automation steps
The dense Apache-2.0 model local-first teams call 'the new baseline'. Excellent at rule-following, OCR and automation; needs ~48GB for full 256K context and is not a frontier reasoner. Free to run, no API price to speak of.
Caveats: Dense 31B: expect ~12-16 tok/s on Apple Silicon at Q4-Q6, and 70GB RAM at full context. Several HN users say Qwen 3.6/3.8 handles long context and agentic tasks a little better; Gemma wins on instruction discipline. AA 15 on the current index; this is a local model, not a hosted-API competitor.
What practitioners say
Gemma 4 31B: HN's local-LLM crowd treats it as the reference point: 'the new baseline for local models' (soganess, 70GB peak on an M5 Max at 256K context); 'particularly good at pipeline/automation tasks' and better than Qwen even at 100B+ for rule-following; 'very good at OCR'. One user had it catch a bug Opus 4.7 missed. Counterpoints: Qwen 3.6 'handles context a little better', peers 'predominantly run Qwen', and ~11-16 tok/s on Macs feels slow.
Questions people ask
Which is better, Gemini 3.5 Flash-Lite or Gemma 4 31B?
Gemini 3.5 Flash-Lite scores higher on the Models at Work leaderboard (63 vs 58 out of 100). On Artificial Analysis' independent Intelligence Index, Gemini 3.5 Flash-Lite leads 22 to 15.
When should I choose Gemini 3.5 Flash-Lite over Gemma 4 31B?
Choose Gemini 3.5 Flash-Lite for classification and extraction pipelines, routing and triage, agentic search at volume. Google's volume tier: 1M context, ~430 tok/s, $0.30/$2.50, and a respectable Arena rank 70. The right default for classification, extraction and routing at scale. Not a reasoning model in any serious sense (AA 22).
When should I choose Gemma 4 31B over Gemini 3.5 Flash-Lite?
Choose Gemma 4 31B for on-device and workstation inference, ocr and document vision, deterministic pipeline/automation steps. The dense Apache-2.0 model local-first teams call 'the new baseline'. Excellent at rule-following, OCR and automation; needs ~48GB for full 256K context and is not a frontier reasoner. Free to run, no API price to speak of.
Scores come from the Models at Work leaderboard: Score is 0 to 100 and computed, not typed: 55% independent evals (Artificial Analysis, Arena, Scale SEAL, Epoch, each scaled against its natural floor and the best score in this table, then averaged), 15% blended price on a fixed log scale ($0.05 per million is 100, $60 is 0), 15% practitioner sentiment from Signals and community threads, 15% operational fit (context window, open weights). A missing component drops out and the rest are reweighted. A model with no independent eval yet is provisional and ranks below every measured one. Within three points is a tie. Vendor-published figures are labelled and never counted. Prices are vendor list prices where published; blended figures assume three input tokens per output token.