Head to head · October 2026

GLM-5.3 vs GPT-6.1 Sol

GLM-5.3 and GPT-6.1 Sol are within three points on the Models at Work leaderboard, which we treat as a tie (81 vs 81). GLM-5.3 is about 1.9× cheaper per blended million tokens ($2.15 vs $4). On Artificial Analysis' independent Intelligence Index, GPT-6.1 Sol leads 52 to 45.

Updated Oct 11, 2026 · every number links to who measured it
GLM-5.3 · Z.aiGPT-6.1 Sol · OpenAI
Leaderboard rank#5#6
Score (0–100)8181
TierOpen weightsWorkhorse
Price per 1M tokens$1.40 in · $4.40 out$2 in · $10 out
Context window1M1M
Open weightsYesNo
Artificial Analysis Intelligence Index4552
Arena Text score1478 (rank 27)—
Scale SWE-Bench Pro V2 (mini-swe-agent)84.30%—
Scale MCP Atlas84.20—
Arena WebDev score1622 (rank 22)—
Humanity's Sixth Sense (Scale)—46.6
Best forAgentic coding in Claude Code / Pi-style harnesses, Security research and code auditing, Self-hosted frontier-ish reasoningOpenAI-standardised teams, Agents API and Decisions API, Cost-controlled reasoning

Choose GLM-5.3 if…

  • you need agentic coding in claude code / pi-style harnesses
  • you need security research and code auditing
  • you need self-hosted frontier-ish reasoning

The open model people actually ship coding agents on. Same 744B base as GLM-5.2 with heavy agentic post-training; drops into Claude Code harnesses and holds its own on SWE-Bench Pro V2 and MCP Atlas at a third of frontier prices.

Caveats: License is muddled: AA lists a custom 'GLM-5.3 License' with commercial restrictions while Arena lists MIT. Read it before you self-host. Permissive on offensive-security tasks; that is a feature for red teams and a governance problem for everyone else. Open weights shipped ~2 weeks after the API; HN worried the public weights were safety-tuned differently.

Choose GPT-6.1 Sol if…

  • you need openai-standardised teams
  • you need agents api and decisions api
  • you need cost-controlled reasoning

Near-Astra intelligence at a fifth of the price, which is OpenAI's own line and, for once, the index agrees.

What practitioners say

GLM-5.3: One of the biggest HN launches of the year (1,171 points, 584 comments). Practitioners report it 'fits Claude Code as its own, zero issues', that '$5 in tokens' of Claude work costs '$0.50' on GLM via Pi, and that it 'routinely finds bugs missed by Fable and Sol' in audits. The cyber angle was the controversy: users bragged about adapting kernel exploits that 'Claude and Opus outright refused', and others flagged the risk of nerfed public weights.

GPT-6.1 Sol: The launch thread's title did the arguing: 'near-Astra for a fifth of the price' drew 955 comments, most of them about routing.

Questions people ask

Which is better, GLM-5.3 or GPT-6.1 Sol?

GLM-5.3 and GPT-6.1 Sol are within three points on the Models at Work leaderboard, which we treat as a tie (81 vs 81). GLM-5.3 is about 1.9× cheaper per blended million tokens ($2.15 vs $4). On Artificial Analysis' independent Intelligence Index, GPT-6.1 Sol leads 52 to 45.

When should I choose GLM-5.3 over GPT-6.1 Sol?

Choose GLM-5.3 for agentic coding in claude code / pi-style harnesses, security research and code auditing, self-hosted frontier-ish reasoning. The open model people actually ship coding agents on. Same 744B base as GLM-5.2 with heavy agentic post-training; drops into Claude Code harnesses and holds its own on SWE-Bench Pro V2 and MCP Atlas at a third of frontier prices.

When should I choose GPT-6.1 Sol over GLM-5.3?

Choose GPT-6.1 Sol for openai-standardised teams, agents api and decisions api, cost-controlled reasoning. Near-Astra intelligence at a fifth of the price, which is OpenAI's own line and, for once, the index agrees.

Which is cheaper, GLM-5.3 or GPT-6.1 Sol?

GLM-5.3 costs $1.40 per million input tokens and $4.40 per million output tokens on Z.ai's list price (Z.ai Pricing, read Oct 11, 2026). GPT-6.1 Sol costs $2 per million input tokens and $10 per million output tokens on OpenAI's list price (OpenAI API pricing, read Oct 11, 2026).

Scores come from the Models at Work leaderboard: Score is 0 to 100 and computed, not typed: 55% independent evals (Artificial Analysis, Arena, Scale SEAL, Epoch, each scaled against its natural floor and the best score in this table, then averaged), 15% blended price on a fixed log scale ($0.05 per million is 100, $60 is 0), 15% practitioner sentiment from Signals and community threads, 15% operational fit (context window, open weights). A missing component drops out and the rest are reweighted. A model with no independent eval yet is provisional and ranks below every measured one. Within three points is a tie. Vendor-published figures are labelled and never counted. Prices are vendor list prices where published; blended figures assume three input tokens per output token.