Leaderboard · #11 of 12 · Moonshot · Open weights

Kimi K3

The highest-placed open-licence model on Arena, and the one most teams have not tried yet.

Updated Oct 10, 2026 · kept by WrenJSON

Key facts

Score
60 / 100, rank #11
Price
$2.31 per 1M blended (Artificial Analysis estimate)
Blended price
$2.31 per 1M tokens (3 input : 1 output), estimated
Context window
Not verified
Open weights
Yes
Best for
Self-hosting, Open-licence requirements, Arena-style chat quality

Moonshot does not publish a simple list price for Kimi K3; Artificial Analysis estimates a blended $2.31 per million tokens (three input to one output).

Independent evals

BenchmarkResultWho ran itRead
Arena text score (max)1488Arena text leaderboard (Oct 8, 2026)Oct 9, 2026
Artificial Analysis Intelligence Index (max)44Artificial Analysis LLM leaderboardOct 10, 2026
  • Artificial Analysis blended price (max): $2.31 per 1M (Artificial Analysis LLM leaderboard)

Caveats

  • Price is Artificial Analysis' blended figure, not a vendor list price; context not verified from a primary source.

Questions people ask

How much does Kimi K3 cost?

Moonshot does not publish a simple list price for Kimi K3; Artificial Analysis estimates a blended $2.31 per million tokens (three input to one output).

How good is Kimi K3?

Kimi K3 ranks #11 of 12 on the Models at Work leaderboard with a score of 60 out of 100 (October 2026 edition). The highest-placed open-licence model on Arena, and the one most teams have not tried yet.

What is Kimi K3 best for?

Kimi K3 is best for self-hosting, open-licence requirements, arena-style chat quality.

What are Kimi K3's benchmark scores?

Arena text score (max): 1488 (Arena text leaderboard (Oct 8, 2026)); Artificial Analysis Intelligence Index (max): 44 (Artificial Analysis LLM leaderboard).

Is Kimi K3 open weights?

Yes. Kimi K3 is released with open weights, so it can be self-hosted.

Compare Kimi K3

Methodology: Score is 0 to 100: 50% independent evals normalised within this table, 20% price-performance, 15% practitioner sentiment from Signals and community threads, 15% operational fit (context, speed, availability, open weights). Models within three points are a tie in practice. Vendor-published figures are labelled as such and never counted as independent.