Leaderboard · #17 of 26 · Alibaba · Open weights

Qwen3.8-27B

The local model r/LocalLLaMA actually converged on this quarter: a dense 27B Apache-2.0 multimodal that fits a single 24GB card at Q4 and does reliable multi-step tool calling. Knowledge recall regressed; it is an agent, not an encyclopedia

Updated Oct 11, 2026 · kept by WrenJSON

Key facts

Score
67 / 100, rank #17
Evals63Price56People76Fit85
Price
$0.50 in · $3 out per 1M tokens (Qwen3.8 27B - Artificial Analysis (median hosted price))
Blended price
$1.13 per 1M tokens (3 input : 1 output)
Context window
256K
Open weights
Yes
Released
Aug 14, 2026
Best for
Local agentic coding on one GPU, On-prem tool-calling agents, Fine-tuning base for vertical assistants

Qwen3.8-27B costs $0.50 per million input tokens and $3 per million output tokens on Alibaba's list price (Qwen3.8 27B - Artificial Analysis (median hosted price), read Oct 11, 2026).

Independent evals

BenchmarkResultWho ran itRead
Artificial Analysis Intelligence Index (xhigh)34Qwen3.8 27B - Artificial AnalysisOct 10, 2026
Arena Text score1438 (rank 99)Text Arena LeaderboardOct 10, 2026
Arena WebDev score1593 (rank 28)WebDev Arena LeaderboardOct 10, 2026

What practitioners say

r/LocalLLaMA's one-week verdict: 'highest level of agency I've ever seen in a local model', with one 3090 running 80 tool calls with zero failures to scrape a university site, and a REST API plus MCP server built from three prompts. Complaints: knowledge regressed vs 3.6, reasoning loops below Q6 quants, and the xhigh default wastes context. HN users report 40+ tok/s on a 4090 at Q4 and ~14 tok/s on a Mac Studio M3 Ultra without MLX tuning.

Caveats

  • Defaults to xhigh reasoning and burns ~39k thinking tokens where low uses ~4k; set medium or lower.
  • Trivia and factual recall are worse than Qwen 3.6 by design.
  • The launch chat template was broken (crashed on enable_thinking=false, bad JSON tool args); use the community-fixed template.

Questions people ask

How much does Qwen3.8-27B cost?

Qwen3.8-27B costs $0.50 per million input tokens and $3 per million output tokens on Alibaba's list price (Qwen3.8 27B - Artificial Analysis (median hosted price), read Oct 11, 2026).

How good is Qwen3.8-27B?

Qwen3.8-27B ranks #17 of 26 on the Models at Work leaderboard with a score of 67 out of 100 (October 2026 edition). The local model r/LocalLLaMA actually converged on this quarter: a dense 27B Apache-2.0 multimodal that fits a single 24GB card at Q4 and does reliable multi-step tool calling. Knowledge recall regressed; it is an agent, not an encyclopedia

What is Qwen3.8-27B best for?

Qwen3.8-27B is best for local agentic coding on one gpu, on-prem tool-calling agents, fine-tuning base for vertical assistants.

What are Qwen3.8-27B's benchmark scores?

Artificial Analysis Intelligence Index (xhigh): 34 (Qwen3.8 27B - Artificial Analysis); Arena Text score: 1438 (rank 99) (Text Arena Leaderboard); Arena WebDev score: 1593 (rank 28) (WebDev Arena Leaderboard).

What is Qwen3.8-27B's context window?

Qwen3.8-27B has a 256K token context window according to Alibaba.

Is Qwen3.8-27B open weights?

Yes. Qwen3.8-27B is released with open weights, so it can be self-hosted.

Compare Qwen3.8-27B

Methodology: Score is 0 to 100 and computed, not typed: 55% independent evals (Artificial Analysis, Arena, Scale SEAL, Epoch, each scaled against its natural floor and the best score in this table, then averaged), 15% blended price on a fixed log scale ($0.05 per million is 100, $60 is 0), 15% practitioner sentiment from Signals and community threads, 15% operational fit (context window, open weights). A missing component drops out and the rest are reweighted. A model with no independent eval yet is provisional and ranks below every measured one. Within three points is a tie. Vendor-published figures are labelled and never counted.