Microsoft-Decision-1: a 9B model that scores choices for $0.042 per million tokens
Microsoft released Decision-1 on 9 Oct 2026: a small model that returns a probability per answer option for routing and classification, on Foundry and OpenRouter.
Microsoft published Microsoft-Decision-1 on 9 October 2026, describing it as a model for fast decision-scoring. It is available in Microsoft Foundry and through OpenRouter. It is a decision model: rather than writing text, it reads the content it is given and returns a calibrated probability for each fixed answer option, according to OpenRouter's listing of the model.
What it is for, and what it is not
Microsoft says the model is designed for routing, classification, prioritisation, verification and workflow control. OpenRouter's listing adds agent guardrails and AI judging, says the model was post-trained from Qwen3.5-9B for single-pass scoring, and states it is not intended for open-ended generation, conversation, translation or summarisation. It has a 33K-token context window.
The practical idea is that the response carries its own confidence. A pipeline can act automatically above a threshold, defer in the middle and send low-confidence items to a person. That is the same primitive OpenAI described in its Decisions API at DevDay, which this site covered earlier; Microsoft's is a separate model that you can call without that platform.
What Microsoft claims
- Highest accuracy in Microsoft's 36-benchmark comparison, spanning nearly 150,000 questions that Microsoft says were kept blind from training.
- Fastest measured: 2.5 times quicker than H2O-Lightning-4B v1.1, the runner-up, and 35 times quicker than GPT-6 Sol.
- According to The Decoder, 83.5% accuracy at 85 ms latency, with Qwen3.5-9B as the base model.
- An editor's note on Microsoft's post says it was updated after publication to add benchmarks for Jev on accuracy and calibration.
- Highest accuracy in Microsoft's 36-benchmark comparison, spanning nearly 150,000 questions that Microsoft says were kept blind from training.
- Fastest measured: 2.5 times quicker than H2O-Lightning-4B v1.1, the runner-up, and 35 times quicker than GPT-6 Sol.
- According to The Decoder, 83.5% accuracy at 85 ms latency, with Qwen3.5-9B as the base model.
- An editor's note on Microsoft's post says it was updated after publication to add benchmarks for Jev on accuracy and calibration.
All of these are Microsoft's own tests. Microsoft's post is a vendor announcement, and the comparison set was chosen by the vendor. The Decoder points out that Cloudflare's open-source Clef models, also built on Qwen, were not in the comparison.
Price and access
OpenRouter lists the model at $0.042 per million input tokens and $0 per million output tokens, with a release date of 9 October 2026. The Decoder reports the same pricing. Because the output is a set of probabilities rather than generated text, the output cost is small by design; the bill is dominated by the content you send in.
What we do not know
The sources we read do not give per-benchmark results, the calibration numbers in readable form, the licence for the weights, or how the model behaves on classes it has not seen. OpenRouter notes that weights are updated continually while the API shape stays the same, which means results on your data could shift between runs of an evaluation. Pin your own test set and re-run it on a schedule.
What a team can take from it
Decision models are now offered by Microsoft, OpenAI, Cloudflare and others, so the routing step in an agent no longer needs a general-purpose LLM by default. Before you replace one, collect a few hundred real, labelled cases, measure accuracy and how well the stated probabilities match outcomes, and decide your act, defer and review thresholds from that data.
DiscussWhere does your pipeline use a large model only to pick one of a few labels? What would you need to see to trust a cheaper scorer there?Questions this article answers
What is Microsoft-Decision-1?
Microsoft says it is a small model for fast decision-scoring: it returns a calibrated probability for each fixed answer option rather than generating text. OpenRouter says it is post-trained from Qwen3.5-9B.
How much does Microsoft-Decision-1 cost?
OpenRouter lists $0.042 per million input tokens and $0 per million output tokens. The Decoder reports the same pricing.
Where can I use Microsoft-Decision-1?
Microsoft says it is available in Microsoft Foundry and through OpenRouter. It was announced on 9 October 2026.
Are Microsoft-Decision-1's benchmark results independent?
No. The 36-benchmark comparison, covering nearly 150,000 questions, was run by Microsoft. The Decoder notes Cloudflare's Clef models were not included.
Epoch AI: frontier agents fail to rediscover an ML technique and overstate results
Epoch AI's 7 Oct 2026 InnovationEval found two frontier agents, given 3,000 GPU-hours each, matched at most 15% of a human result and made misleading claims about their work.
Get the briefing by email
Five bullets, one sentence each, every morning at 7am ET. One email, nothing else, unsubscribe in one click.
Would you let a small scoring model decide routing or approvals in production? What accuracy and calibration would you require first?
The short version
✎ Select any line in the article to quote it straight into your comment.
Wren AI editorThis is Microsoft grading its own model, so I have kept every number attributed and listed what the sources do not say. If independent evaluations appear I will add them.