Head to head · October 2026

Qwen3.8 Max vs Qwen3.8-27B

Qwen3.8 Max scores higher on the Models at Work leaderboard (78 vs 67 out of 100). Qwen3.8-27B is about 2.7× cheaper per blended million tokens ($1.13 vs $3). On Artificial Analysis' independent Intelligence Index, Qwen3.8 Max leads 45 to 34.

Updated Oct 11, 2026 · every number links to who measured it
Qwen3.8 Max · AlibabaQwen3.8-27B · Alibaba
Leaderboard rank#11#17
Score (0–100)7867
TierWorkhorseOpen weights
Price per 1M tokens$2 in · $6 out$0.50 in · $3 out
Context window984K256K
Open weightsNoYes
Artificial Analysis Intelligence Index4534
Arena Text score1483 (rank 22)1438 (rank 99)
Arena WebDev score (qwen3.8-max-0902)1674 (rank 10)—
Scale MCP Atlas (Qwen3.8-2.4T-A95B open-weight variant, xHigh)84.50—
Arena WebDev score—1593 (rank 28)
Best forAgentic coding, Tool/MCP-heavy agents, Multilingual production chatLocal agentic coding on one GPU, On-prem tool-calling agents, Fine-tuning base for vertical assistants

Choose Qwen3.8 Max if…

  • you need agentic coding
  • you need tool/mcp-heavy agents
  • you need multilingual production chat

Alibaba's hosted flagship: a 2.4T MoE that sits just under the Western frontier and above every other open-lab model on Arena WebDev. Priced like a mid-tier US model, not like a Chinese one. Strong at coding and MCP tool use.

Caveats: At $2/$6 it costs 4-5x GLM-5.3 or MiMo-V2.6-Pro for a similar AA score. The hosted Max (0902) is proprietary; the open-weight sibling is Qwen3.8-2.4T-A95B, which scores lower on AA (40) and needs serious hardware. Slow: ~35 tok/s on AA's measurement.

Choose Qwen3.8-27B if…

  • you need local agentic coding on one gpu
  • you need on-prem tool-calling agents
  • you need fine-tuning base for vertical assistants

The local model r/LocalLLaMA actually converged on this quarter: a dense 27B Apache-2.0 multimodal that fits a single 24GB card at Q4 and does reliable multi-step tool calling. Knowledge recall regressed; it is an agent, not an encyclopedia

Caveats: Defaults to xhigh reasoning and burns ~39k thinking tokens where low uses ~4k; set medium or lower. Trivia and factual recall are worse than Qwen 3.6 by design. The launch chat template was broken (crashed on enable_thinking=false, bad JSON tool args); use the community-fixed template.

What practitioners say

Qwen3.8 Max: The launch thread (1,124 points, 612 comments) was dominated by people already running the smaller Qwen 3.6/3.8 line locally and cancelling Claude subscriptions; the Max itself drew praise as the first Qwen-Max-class model with open weights. Skeptics (Aurornis) say Qwen output 'has to be discarded for anything other than really easy tasks', and several note that at current electricity prices self-hosting is pricier than DeepSeek's cache rate.

Qwen3.8-27B: r/LocalLLaMA's one-week verdict: 'highest level of agency I've ever seen in a local model', with one 3090 running 80 tool calls with zero failures to scrape a university site, and a REST API plus MCP server built from three prompts. Complaints: knowledge regressed vs 3.6, reasoning loops below Q6 quants, and the xhigh default wastes context. HN users report 40+ tok/s on a 4090 at Q4 and ~14 tok/s on a Mac Studio M3 Ultra without MLX tuning.

Questions people ask

Which is better, Qwen3.8 Max or Qwen3.8-27B?

Qwen3.8 Max scores higher on the Models at Work leaderboard (78 vs 67 out of 100). Qwen3.8-27B is about 2.7× cheaper per blended million tokens ($1.13 vs $3). On Artificial Analysis' independent Intelligence Index, Qwen3.8 Max leads 45 to 34.

When should I choose Qwen3.8 Max over Qwen3.8-27B?

Choose Qwen3.8 Max for agentic coding, tool/mcp-heavy agents, multilingual production chat. Alibaba's hosted flagship: a 2.4T MoE that sits just under the Western frontier and above every other open-lab model on Arena WebDev. Priced like a mid-tier US model, not like a Chinese one. Strong at coding and MCP tool use.

When should I choose Qwen3.8-27B over Qwen3.8 Max?

Choose Qwen3.8-27B for local agentic coding on one gpu, on-prem tool-calling agents, fine-tuning base for vertical assistants. The local model r/LocalLLaMA actually converged on this quarter: a dense 27B Apache-2.0 multimodal that fits a single 24GB card at Q4 and does reliable multi-step tool calling. Knowledge recall regressed; it is an agent, not an encyclopedia

Which is cheaper, Qwen3.8 Max or Qwen3.8-27B?

Qwen3.8 Max costs $2 per million input tokens and $6 per million output tokens on Alibaba's list price (Qwen3.8 Max (0902) - Artificial Analysis, read Oct 11, 2026). Qwen3.8-27B costs $0.50 per million input tokens and $3 per million output tokens on Alibaba's list price (Qwen3.8 27B - Artificial Analysis (median hosted price), read Oct 11, 2026).

Scores come from the Models at Work leaderboard: Score is 0 to 100 and computed, not typed: 55% independent evals (Artificial Analysis, Arena, Scale SEAL, Epoch, each scaled against its natural floor and the best score in this table, then averaged), 15% blended price on a fixed log scale ($0.05 per million is 100, $60 is 0), 15% practitioner sentiment from Signals and community threads, 15% operational fit (context window, open weights). A missing component drops out and the rest are reweighted. A model with no independent eval yet is provisional and ranks below every measured one. Within three points is a tie. Vendor-published figures are labelled and never counted. Prices are vendor list prices where published; blended figures assume three input tokens per output token.