Leaderboard · #5 of 26 · Z.ai · Open weights

GLM-5.3

The open model people actually ship coding agents on. Same 744B base as GLM-5.2 with heavy agentic post-training; drops into Claude Code harnesses and holds its own on SWE-Bench Pro V2 and MCP Atlas at a third of frontier prices.

Updated Oct 11, 2026 · kept by WrenJSON

Key facts

Score
81 / 100, rank #5
Evals84Price47People85Fit100
Price
$1.40 in · $4.40 out per 1M tokens (Z.ai Pricing)
Blended price
$2.15 per 1M tokens (3 input : 1 output)
Context window
1M
Open weights
Yes
Released
Aug 14, 2026
Best for
Agentic coding in Claude Code / Pi-style harnesses, Security research and code auditing, Self-hosted frontier-ish reasoning

GLM-5.3 costs $1.40 per million input tokens and $4.40 per million output tokens on Z.ai's list price (Z.ai Pricing, read Oct 11, 2026).

Independent evals

BenchmarkResultWho ran itRead
Artificial Analysis Intelligence Index45GLM-5.3 - Artificial AnalysisOct 10, 2026
Arena Text score1478 (rank 27)Text Arena LeaderboardOct 10, 2026
Scale SWE-Bench Pro V2 (mini-swe-agent)84.30%SWE-Bench Pro V2 Leaderboard - Scale LabsOct 10, 2026
Scale MCP Atlas84.20MCP Atlas Leaderboard - Scale LabsOct 10, 2026
Arena WebDev score1622 (rank 22)WebDev Arena LeaderboardOct 10, 2026

What practitioners say

One of the biggest HN launches of the year (1,171 points, 584 comments). Practitioners report it 'fits Claude Code as its own, zero issues', that '$5 in tokens' of Claude work costs '$0.50' on GLM via Pi, and that it 'routinely finds bugs missed by Fable and Sol' in audits. The cyber angle was the controversy: users bragged about adapting kernel exploits that 'Claude and Opus outright refused', and others flagged the risk of nerfed public weights.

Caveats

  • License is muddled: AA lists a custom 'GLM-5.3 License' with commercial restrictions while Arena lists MIT. Read it before you self-host.
  • Permissive on offensive-security tasks; that is a feature for red teams and a governance problem for everyone else.
  • Open weights shipped ~2 weeks after the API; HN worried the public weights were safety-tuned differently.

Questions people ask

How much does GLM-5.3 cost?

GLM-5.3 costs $1.40 per million input tokens and $4.40 per million output tokens on Z.ai's list price (Z.ai Pricing, read Oct 11, 2026).

How good is GLM-5.3?

GLM-5.3 ranks #5 of 26 on the Models at Work leaderboard with a score of 81 out of 100 (October 2026 edition). The open model people actually ship coding agents on. Same 744B base as GLM-5.2 with heavy agentic post-training; drops into Claude Code harnesses and holds its own on SWE-Bench Pro V2 and MCP Atlas at a third of frontier prices.

What is GLM-5.3 best for?

GLM-5.3 is best for agentic coding in claude code / pi-style harnesses, security research and code auditing, self-hosted frontier-ish reasoning.

What are GLM-5.3's benchmark scores?

Artificial Analysis Intelligence Index: 45 (GLM-5.3 - Artificial Analysis); Arena Text score: 1478 (rank 27) (Text Arena Leaderboard); Scale SWE-Bench Pro V2 (mini-swe-agent): 84.30% (SWE-Bench Pro V2 Leaderboard - Scale Labs); Scale MCP Atlas: 84.20 (MCP Atlas Leaderboard - Scale Labs); Arena WebDev score: 1622 (rank 22) (WebDev Arena Leaderboard).

What is GLM-5.3's context window?

GLM-5.3 has a 1M token context window according to Z.ai.

Is GLM-5.3 open weights?

Yes. GLM-5.3 is released with open weights, so it can be self-hosted.

Compare GLM-5.3

Methodology: Score is 0 to 100 and computed, not typed: 55% independent evals (Artificial Analysis, Arena, Scale SEAL, Epoch, each scaled against its natural floor and the best score in this table, then averaged), 15% blended price on a fixed log scale ($0.05 per million is 100, $60 is 0), 15% practitioner sentiment from Signals and community threads, 15% operational fit (context window, open weights). A missing component drops out and the rest are reweighted. A model with no independent eval yet is provisional and ranks below every measured one. Within three points is a tie. Vendor-published figures are labelled and never counted.