Nemotron 3 Ultra
The strongest US-made open model, built for throughput: a 550B hybrid Mamba-Transformer MoE that NVIDIA tunes to run fast on its own hardware. Trails the Chinese open models on intelligence and has lost ground since June.
Key facts
- Score
- 60 / 100, rank #21Evals57Price59People45Fit85
- Price
- $0.50 in · $2.20 out per 1M tokens (Nemotron 3 Ultra - OpenRouter)
- Blended price
- $0.93 per 1M tokens (3 input : 1 output)
- Context window
- 262K
- Open weights
- Yes
- Released
- Jun 4, 2026
- Best for
- High-throughput agent pipelines on NVIDIA infra, Teams that need a non-Chinese open-weight model, Enterprise RAG and orchestration
Nemotron 3 Ultra costs $0.50 per million input tokens and $2.20 per million output tokens on NVIDIA's list price (Nemotron 3 Ultra - OpenRouter, read Oct 11, 2026).
Independent evals
| Benchmark | Result | Who ran it | Read |
|---|---|---|---|
| Artificial Analysis Intelligence Index | 23 | LLM Leaderboard - Artificial Analysis | Oct 10, 2026 |
| Arena Text score (nvidia-nemotron-3-ultra-550b) | 1426 (rank 120) | Text Arena Leaderboard | Oct 10, 2026 |
| Scale MCP Atlas (thinking) | 63.10 | MCP Atlas Leaderboard - Scale Labs | Oct 10, 2026 |
- Artificial Analysis output speed (tokens/s): 142 (LLM Leaderboard - Artificial Analysis)
What practitioners say
Little organic practitioner chatter compared with the Chinese open models. Artificial Analysis's launch note called it 'leading US open weights intelligence' at 'over 300 tokens per second' but still behind Kimi K2.6. A WebBrain planner benchmark via OpenRouter scored it 81/100 but found a 40.6s p95 latency 'hard to imagine' for interactive use. Most discussion centres on the architecture (LatentMoE, multi-token prediction) rather than day-to-day coding results.
Caveats
- AA scored it 48 on the launch-era index; on the current v4.3.2 index it sits at 23, well behind GLM-5.3, MiMo and Qwen3.8.
- Hosted on OpenRouter at $0.50/$2.20 (plus a free tier); NVIDIA itself does not publish a per-token price.
- 262K context on BF16 weights despite 1M marketing in some hosts; one planner benchmark saw a 40.6s p95 latency.
Questions people ask
How much does Nemotron 3 Ultra cost?
Nemotron 3 Ultra costs $0.50 per million input tokens and $2.20 per million output tokens on NVIDIA's list price (Nemotron 3 Ultra - OpenRouter, read Oct 11, 2026).
How good is Nemotron 3 Ultra?
Nemotron 3 Ultra ranks #21 of 26 on the Models at Work leaderboard with a score of 60 out of 100 (October 2026 edition). The strongest US-made open model, built for throughput: a 550B hybrid Mamba-Transformer MoE that NVIDIA tunes to run fast on its own hardware. Trails the Chinese open models on intelligence and has lost ground since June.
What is Nemotron 3 Ultra best for?
Nemotron 3 Ultra is best for high-throughput agent pipelines on nvidia infra, teams that need a non-chinese open-weight model, enterprise rag and orchestration.
What are Nemotron 3 Ultra's benchmark scores?
Artificial Analysis Intelligence Index: 23 (LLM Leaderboard - Artificial Analysis); Arena Text score (nvidia-nemotron-3-ultra-550b): 1426 (rank 120) (Text Arena Leaderboard); Scale MCP Atlas (thinking): 63.10 (MCP Atlas Leaderboard - Scale Labs).
What is Nemotron 3 Ultra's context window?
Nemotron 3 Ultra has a 262K token context window according to NVIDIA.
Is Nemotron 3 Ultra open weights?
Yes. Nemotron 3 Ultra is released with open weights, so it can be self-hosted.
Compare Nemotron 3 Ultra
Methodology: Score is 0 to 100 and computed, not typed: 55% independent evals (Artificial Analysis, Arena, Scale SEAL, Epoch, each scaled against its natural floor and the best score in this table, then averaged), 15% blended price on a fixed log scale ($0.05 per million is 100, $60 is 0), 15% practitioner sentiment from Signals and community threads, 15% operational fit (context window, open weights). A missing component drops out and the rest are reweighted. A model with no independent eval yet is provisional and ranks below every measured one. Within three points is a tie. Vendor-published figures are labelled and never counted.