Leaderboard · #9 of 26 · Z.ai · Open weights

GLM-5.3-Flash

Probably the best price-to-intelligence ratio on the market: AA 42 for $0.15/$0.50, MIT-licensed, natively multimodal, 1M context. Not actually fast despite the name; treat it as a cheap executor behind a stronger planner.

Updated Oct 11, 2026 · kept by WrenJSON

Key facts

Score
79 / 100, rank #9
Evals75Price78People74Fit100
Price
$0.15 in · $0.50 out per 1M tokens (Z.ai Pricing)
Blended price
$0.24 per 1M tokens (3 input : 1 output)
Context window
1M
Open weights
Yes
Released
Aug 26, 2026
Best for
Cheap agentic coding execution, Bulk document and image processing, Self-hosted multimodal assistants

GLM-5.3-Flash costs $0.15 per million input tokens and $0.50 per million output tokens on Z.ai's list price (Z.ai Pricing, read Oct 11, 2026).

Independent evals

BenchmarkResultWho ran itRead
Artificial Analysis Intelligence Index42GLM-5.3-Flash - Artificial AnalysisOct 10, 2026
Arena Text score1475 (rank 38)Text Arena LeaderboardOct 10, 2026
Arena WebDev score1609 (rank 26)WebDev Arena LeaderboardOct 10, 2026
  • Artificial Analysis output speed (tokens/s): 59.8 (GLM-5.3-Flash - Artificial Analysis)

What practitioners say

A month-long HN diary ('One month coding with GLM 5.3 Flash', 233 points) spent $68 total; commenters agreed 'a month of agentic coding for $68 is the headline'. The working pattern is 'GLM-5.3 to write the plan, Flash to implement', with one user noting that given an unambiguous plan 'GLM 5.3 Flash executes it just fine'. Negatives: 'too slow for execution, despite the name' on Z.ai hosting, and the full model is 'way more consistent'.

Caveats

  • 'Too slow for execution, despite the name' on Z.ai's own hosting (~60 tok/s); third-party hosts vary.
  • Users report less consistent output than full GLM-5.3; pair it with a planning model.
  • 320B total params: self-hosting still needs a multi-GPU node.

Questions people ask

How much does GLM-5.3-Flash cost?

GLM-5.3-Flash costs $0.15 per million input tokens and $0.50 per million output tokens on Z.ai's list price (Z.ai Pricing, read Oct 11, 2026).

How good is GLM-5.3-Flash?

GLM-5.3-Flash ranks #9 of 26 on the Models at Work leaderboard with a score of 79 out of 100 (October 2026 edition). Probably the best price-to-intelligence ratio on the market: AA 42 for $0.15/$0.50, MIT-licensed, natively multimodal, 1M context. Not actually fast despite the name; treat it as a cheap executor behind a stronger planner.

What is GLM-5.3-Flash best for?

GLM-5.3-Flash is best for cheap agentic coding execution, bulk document and image processing, self-hosted multimodal assistants.

What are GLM-5.3-Flash's benchmark scores?

Artificial Analysis Intelligence Index: 42 (GLM-5.3-Flash - Artificial Analysis); Arena Text score: 1475 (rank 38) (Text Arena Leaderboard); Arena WebDev score: 1609 (rank 26) (WebDev Arena Leaderboard).

What is GLM-5.3-Flash's context window?

GLM-5.3-Flash has a 1M token context window according to Z.ai.

Is GLM-5.3-Flash open weights?

Yes. GLM-5.3-Flash is released with open weights, so it can be self-hosted.

Compare GLM-5.3-Flash

Methodology: Score is 0 to 100 and computed, not typed: 55% independent evals (Artificial Analysis, Arena, Scale SEAL, Epoch, each scaled against its natural floor and the best score in this table, then averaged), 15% blended price on a fixed log scale ($0.05 per million is 100, $60 is 0), 15% practitioner sentiment from Signals and community threads, 15% operational fit (context window, open weights). A missing component drops out and the rest are reweighted. A model with no independent eval yet is provisional and ranks below every measured one. Within three points is a tie. Vendor-published figures are labelled and never counted.