Qwen3.8 Max vs Qwen3.8-27B
Qwen3.8 Max scores higher on the Models at Work leaderboard (78 vs 67 out of 100). Qwen3.8-27B is about 2.7× cheaper per blended million tokens ($1.13 vs $3). On Artificial Analysis' independent Intelligence Index, Qwen3.8 Max leads 45 to 34.
| Qwen3.8 Max · Alibaba | Qwen3.8-27B · Alibaba | |
|---|---|---|
| Leaderboard rank | #11 | #17 |
| Score (0–100) | 78 | 67 |
| Tier | Workhorse | Open weights |
| Price per 1M tokens | $2 in · $6 out | $0.50 in · $3 out |
| Context window | 984K | 256K |
| Open weights | No | Yes |
| Artificial Analysis Intelligence Index | 45 | 34 |
| Arena Text score | 1483 (rank 22) | 1438 (rank 99) |
| Arena WebDev score (qwen3.8-max-0902) | 1674 (rank 10) | — |
| Scale MCP Atlas (Qwen3.8-2.4T-A95B open-weight variant, xHigh) | 84.50 | — |
| Arena WebDev score | — | 1593 (rank 28) |
| Best for | Agentic coding, Tool/MCP-heavy agents, Multilingual production chat | Local agentic coding on one GPU, On-prem tool-calling agents, Fine-tuning base for vertical assistants |
Choose Qwen3.8 Max if…
- you need agentic coding
- you need tool/mcp-heavy agents
- you need multilingual production chat
Alibaba's hosted flagship: a 2.4T MoE that sits just under the Western frontier and above every other open-lab model on Arena WebDev. Priced like a mid-tier US model, not like a Chinese one. Strong at coding and MCP tool use.
Caveats: At $2/$6 it costs 4-5x GLM-5.3 or MiMo-V2.6-Pro for a similar AA score. The hosted Max (0902) is proprietary; the open-weight sibling is Qwen3.8-2.4T-A95B, which scores lower on AA (40) and needs serious hardware. Slow: ~35 tok/s on AA's measurement.
Choose Qwen3.8-27B if…
- you need local agentic coding on one gpu
- you need on-prem tool-calling agents
- you need fine-tuning base for vertical assistants
The local model r/LocalLLaMA actually converged on this quarter: a dense 27B Apache-2.0 multimodal that fits a single 24GB card at Q4 and does reliable multi-step tool calling. Knowledge recall regressed; it is an agent, not an encyclopedia
Caveats: Defaults to xhigh reasoning and burns ~39k thinking tokens where low uses ~4k; set medium or lower. Trivia and factual recall are worse than Qwen 3.6 by design. The launch chat template was broken (crashed on enable_thinking=false, bad JSON tool args); use the community-fixed template.
What practitioners say
Qwen3.8 Max: The launch thread (1,124 points, 612 comments) was dominated by people already running the smaller Qwen 3.6/3.8 line locally and cancelling Claude subscriptions; the Max itself drew praise as the first Qwen-Max-class model with open weights. Skeptics (Aurornis) say Qwen output 'has to be discarded for anything other than really easy tasks', and several note that at current electricity prices self-hosting is pricier than DeepSeek's cache rate.
Qwen3.8-27B: r/LocalLLaMA's one-week verdict: 'highest level of agency I've ever seen in a local model', with one 3090 running 80 tool calls with zero failures to scrape a university site, and a REST API plus MCP server built from three prompts. Complaints: knowledge regressed vs 3.6, reasoning loops below Q6 quants, and the xhigh default wastes context. HN users report 40+ tok/s on a 4090 at Q4 and ~14 tok/s on a Mac Studio M3 Ultra without MLX tuning.
Questions people ask
Which is better, Qwen3.8 Max or Qwen3.8-27B?
Qwen3.8 Max scores higher on the Models at Work leaderboard (78 vs 67 out of 100). Qwen3.8-27B is about 2.7× cheaper per blended million tokens ($1.13 vs $3). On Artificial Analysis' independent Intelligence Index, Qwen3.8 Max leads 45 to 34.
When should I choose Qwen3.8 Max over Qwen3.8-27B?
Choose Qwen3.8 Max for agentic coding, tool/mcp-heavy agents, multilingual production chat. Alibaba's hosted flagship: a 2.4T MoE that sits just under the Western frontier and above every other open-lab model on Arena WebDev. Priced like a mid-tier US model, not like a Chinese one. Strong at coding and MCP tool use.
When should I choose Qwen3.8-27B over Qwen3.8 Max?
Choose Qwen3.8-27B for local agentic coding on one gpu, on-prem tool-calling agents, fine-tuning base for vertical assistants. The local model r/LocalLLaMA actually converged on this quarter: a dense 27B Apache-2.0 multimodal that fits a single 24GB card at Q4 and does reliable multi-step tool calling. Knowledge recall regressed; it is an agent, not an encyclopedia
Which is cheaper, Qwen3.8 Max or Qwen3.8-27B?
Qwen3.8 Max costs $2 per million input tokens and $6 per million output tokens on Alibaba's list price (Qwen3.8 Max (0902) - Artificial Analysis, read Oct 11, 2026). Qwen3.8-27B costs $0.50 per million input tokens and $3 per million output tokens on Alibaba's list price (Qwen3.8 27B - Artificial Analysis (median hosted price), read Oct 11, 2026).
Scores come from the Models at Work leaderboard: Score is 0 to 100 and computed, not typed: 55% independent evals (Artificial Analysis, Arena, Scale SEAL, Epoch, each scaled against its natural floor and the best score in this table, then averaged), 15% blended price on a fixed log scale ($0.05 per million is 100, $60 is 0), 15% practitioner sentiment from Signals and community threads, 15% operational fit (context window, open weights). A missing component drops out and the rest are reweighted. A model with no independent eval yet is provisional and ranks below every measured one. Within three points is a tie. Vendor-published figures are labelled and never counted. Prices are vendor list prices where published; blended figures assume three input tokens per output token.