Gemini 3.8 Flash vs Gemma 4 31B
Gemini 3.8 Flash scores higher on the Models at Work leaderboard (77 vs 58 out of 100). On Artificial Analysis' independent Intelligence Index, Gemini 3.8 Flash leads 41 to 15.
| Gemini 3.8 Flash · Google | Gemma 4 31B · Google | |
|---|---|---|
| Leaderboard rank | #12 | #22 |
| Score (0–100) | 77 | 58 |
| Tier | Fast and cheap | Open weights |
| Price per 1M tokens | $0.75 in · $3.75 out | Not verified |
| Context window | 1M | 256K |
| Open weights | No | Yes |
| Artificial Analysis Intelligence Index | 41 | 15 |
| Arena Text score | 1497 (rank 8) | 1452 (rank 75) |
| Scale SWE-Bench Pro V2 (mini-swe-agent) | 58.80% | — |
| Best for | Agentic coding on a budget, Frontend generation, High-volume multimodal pipelines | On-device and workstation inference, OCR and document vision, Deterministic pipeline/automation steps |
Choose Gemini 3.8 Flash if…
- you need agentic coding on a budget
- you need frontend generation
- you need high-volume multimodal pipelines
The best cheap model for long-running agent loops: rank 8 on Arena, 128 tok/s, and the only sub-$1 model on Scale's SWE-Bench Pro V2 board. Hallucinates more than Claude-class models and the price doubles in January.
Caveats: $0.75/$3.75 is introductory; Google's pricing page says it goes to $1.50/$7.50 on 2027-01-01. HN users report more hallucination and dropped context between turns than rival models; verify outputs. Third Flash release in six weeks (3.6, 3.7, 3.8). Expect short model lifecycles and deprecation churn.
Choose Gemma 4 31B if…
- you need on-device and workstation inference
- you need ocr and document vision
- you need deterministic pipeline/automation steps
The dense Apache-2.0 model local-first teams call 'the new baseline'. Excellent at rule-following, OCR and automation; needs ~48GB for full 256K context and is not a frontier reasoner. Free to run, no API price to speak of.
Caveats: Dense 31B: expect ~12-16 tok/s on Apple Silicon at Q4-Q6, and 70GB RAM at full context. Several HN users say Qwen 3.6/3.8 handles long context and agentic tasks a little better; Gemma wins on instruction discipline. AA 15 on the current index; this is a local model, not a hosted-API competitor.
What practitioners say
Gemini 3.8 Flash: Huge HN thread (1,160 points, 669 comments). simonw built an HTML visualisation in '13 seconds' for '1.8 cents'; colechristensen calls it 'competitive with opus/fable and also FAST'; several use it as the workhorse model with a stronger planner. The pushback is about reliability: one user says it 'gives me the most hallucinations' of the major models, others complain it ignores context between consecutive messages, and the uneven knowledge cutoff (some domains stuck at early 2025) bites in research tasks.
Gemma 4 31B: HN's local-LLM crowd treats it as the reference point: 'the new baseline for local models' (soganess, 70GB peak on an M5 Max at 256K context); 'particularly good at pipeline/automation tasks' and better than Qwen even at 100B+ for rule-following; 'very good at OCR'. One user had it catch a bug Opus 4.7 missed. Counterpoints: Qwen 3.6 'handles context a little better', peers 'predominantly run Qwen', and ~11-16 tok/s on Macs feels slow.
Questions people ask
Which is better, Gemini 3.8 Flash or Gemma 4 31B?
Gemini 3.8 Flash scores higher on the Models at Work leaderboard (77 vs 58 out of 100). On Artificial Analysis' independent Intelligence Index, Gemini 3.8 Flash leads 41 to 15.
When should I choose Gemini 3.8 Flash over Gemma 4 31B?
Choose Gemini 3.8 Flash for agentic coding on a budget, frontend generation, high-volume multimodal pipelines. The best cheap model for long-running agent loops: rank 8 on Arena, 128 tok/s, and the only sub-$1 model on Scale's SWE-Bench Pro V2 board. Hallucinates more than Claude-class models and the price doubles in January.
When should I choose Gemma 4 31B over Gemini 3.8 Flash?
Choose Gemma 4 31B for on-device and workstation inference, ocr and document vision, deterministic pipeline/automation steps. The dense Apache-2.0 model local-first teams call 'the new baseline'. Excellent at rule-following, OCR and automation; needs ~48GB for full 256K context and is not a frontier reasoner. Free to run, no API price to speak of.
Scores come from the Models at Work leaderboard: Score is 0 to 100 and computed, not typed: 55% independent evals (Artificial Analysis, Arena, Scale SEAL, Epoch, each scaled against its natural floor and the best score in this table, then averaged), 15% blended price on a fixed log scale ($0.05 per million is 100, $60 is 0), 15% practitioner sentiment from Signals and community threads, 15% operational fit (context window, open weights). A missing component drops out and the rest are reweighted. A model with no independent eval yet is provisional and ranks below every measured one. Within three points is a tie. Vendor-published figures are labelled and never counted. Prices are vendor list prices where published; blended figures assume three input tokens per output token.