Microsoft published Microsoft-Decision-1 on 9 October 2026, describing it as a model for fast decision-scoring. It is available in Microsoft Foundry and through OpenRouter. It is a decision model: rather than writing text, it reads the content it is given and returns a calibrated probability for each fixed answer option, according to OpenRouter's listing of the model.

What it is for, and what it is not

Microsoft says the model is designed for routing, classification, prioritisation, verification and workflow control. OpenRouter's listing adds agent guardrails and AI judging, says the model was post-trained from Qwen3.5-9B for single-pass scoring, and states it is not intended for open-ended generation, conversation, translation or summarisation. It has a 33K-token context window.

The practical idea is that the response carries its own confidence. A pipeline can act automatically above a threshold, defer in the middle and send low-confidence items to a person. That is the same primitive OpenAI described in its Decisions API at DevDay, which this site covered earlier; Microsoft's is a separate model that you can call without that platform.

What Microsoft claims

  • Highest accuracy in Microsoft's 36-benchmark comparison, spanning nearly 150,000 questions that Microsoft says were kept blind from training.
  • Fastest measured: 2.5 times quicker than H2O-Lightning-4B v1.1, the runner-up, and 35 times quicker than GPT-6 Sol.
  • According to The Decoder, 83.5% accuracy at 85 ms latency, with Qwen3.5-9B as the base model.
  • An editor's note on Microsoft's post says it was updated after publication to add benchmarks for Jev on accuracy and calibration.
  • Highest accuracy in Microsoft's 36-benchmark comparison, spanning nearly 150,000 questions that Microsoft says were kept blind from training.
  • Fastest measured: 2.5 times quicker than H2O-Lightning-4B v1.1, the runner-up, and 35 times quicker than GPT-6 Sol.
  • According to The Decoder, 83.5% accuracy at 85 ms latency, with Qwen3.5-9B as the base model.
  • An editor's note on Microsoft's post says it was updated after publication to add benchmarks for Jev on accuracy and calibration.

All of these are Microsoft's own tests. Microsoft's post is a vendor announcement, and the comparison set was chosen by the vendor. The Decoder points out that Cloudflare's open-source Clef models, also built on Qwen, were not in the comparison.

Price and access

OpenRouter lists the model at $0.042 per million input tokens and $0 per million output tokens, with a release date of 9 October 2026. The Decoder reports the same pricing. Because the output is a set of probabilities rather than generated text, the output cost is small by design; the bill is dominated by the content you send in.

What we do not know

The sources we read do not give per-benchmark results, the calibration numbers in readable form, the licence for the weights, or how the model behaves on classes it has not seen. OpenRouter notes that weights are updated continually while the API shape stays the same, which means results on your data could shift between runs of an evaluation. Pin your own test set and re-run it on a schedule.

What a team can take from it

Decision models are now offered by Microsoft, OpenAI, Cloudflare and others, so the routing step in an agent no longer needs a general-purpose LLM by default. Before you replace one, collect a few hundred real, labelled cases, measure accuracy and how well the stated probabilities match outcomes, and decide your act, defer and review thresholds from that data.

DiscussWhere does your pipeline use a large model only to pick one of a few labels? What would you need to see to trust a cheaper scorer there?