block #0009 in --:--:--Join the pool
PENDING…
ai ✓ confirmed 6/6 1h ago · 3 min read

Microsoft-Decision-1 model Scores Choices 35x Faster Than GPT-6 Sol

Microsoft-Decision-1 model returns a calibrated probability for each fixed option, costs $0.042 per million input tokens with free output, and claims 35x the speed of GPT-6 Sol.

Microsoft-Decision-1 model Scores Choices 35x Faster Than GPT-6 Sol
tl;dr
  • Microsoft-Decision-1 is post-trained from Qwen3.5-9B and handles yes/no, multiple-choice, rating and rubric-grading decisions in a single pass.
  • Pricing is $0.042 per million input tokens, and output tokens are free, available in Microsoft Foundry.
  • Microsoft says it topped accuracy in its own 36-benchmark comparison with nearly 150,000 questions kept blind from training.
in this block
  1. What actually happened
  2. The speed and quality claims
  3. Inside Microsoft
  4. Availability, with one conflict
  5. Why decision models are suddenly hot
  6. What to do as a reader

The Microsoft-Decision-1 model is Microsoft's entry into the "decision model" craze, and the pitch is simple: stop paying a big LLM to answer yes or no. Announced on October 9, it reads your content, looks at a fixed list of options and returns a calibrated probability for each one. Microsoft claims it is 35 times faster than GPT-6 Sol, with output tokens priced at zero.

What actually happened

In a Microsoft Command Line post dated October 9, Achint Srivastava, VP of Software Engineering in Microsoft's Office of the CTO, introduced the model for routing, classification, prioritization, verification and workflow control. Unlike chat models, it doesn't write free text. It scores the options you give it. Srivastava calls decision models "an important new category in AI," built to deliver structured outputs that software can act on immediately.

The OpenRouter listing describes the target jobs as classification, routing, prioritization, verification, workflow control, agent guardrails and AI judging. It also spells out what the model is not for: open-ended generation, conversation, translation or summarization.

Microsoft says it post-trained Qwen3.5-9B for fast, single-pass scoring and will soon rebase the model on others, including Microsoft AI (MAI) and OpenAI models. The OpenRouter listing adds that weights are updated continually while the API shape stays the same, which matters if you need reproducible behavior.

The speed and quality claims

Microsoft says the Microsoft-Decision-1 model was the fastest it measured: 2.5 times quicker than runner-up H2O-Lightning-4B v1.1 and 35 times quicker than GPT-6 Sol. AlphaSignal's writeup instead cites a P50 comparison of 35 times faster than GPT-6 Sol and 4.5 times faster than Quyet-1.0-Large. Different baselines, same headline, so we're listing both.

On robustness, Microsoft perturbed requests in eight ways, like reordering options or adding harmless formatting noise. The model changed its decision on 1.3% of perturbations on average, with zero flips when option descriptions were paraphrased, reversed or shuffled. Microsoft also notes the post was updated to add benchmarks for Jev on accuracy and calibration.

All of this is vendor-reported. AlphaSignal points out that several comparisons omit full baseline details, and that production results depend on your prompts, labels and thresholds.

Inside Microsoft

Microsoft shared four internal uses:

  • Xbox researchers sorted survey answers and reviews from Steam and X into themes, with quality competitive with GPT-6 Sol at over 14 times the speed and 200 times lower cost.
  • The Copilot team found it competitive with GPT5.6 Luna for quality checks while 100 times faster.
  • On-call engineers use it to pick relevant knowledge during incidents.
  • A research team scored adaptive replanning 46 times more consistently than an LLM-based score at three times the speed.

Availability, with one conflict

Microsoft says the model is available in Foundry and through OpenRouter, and the OpenRouter page is live with one provider. AlphaSignal, however, wrote that OpenRouter support was "planned." Check OpenRouter directly before you build around it.

The Microsoft-Decision-1 model joins a crowded field. We've covered the OpenAI Decisions API, Liquid AI's D1 and the Jev 27B decision model, all chasing the same job: cheap, fast verdicts inside agent pipelines.

Why decision models are suddenly hot

Agents make dozens of small calls: should this step continue, does this output pass, which tool fits. AlphaSignal's example is blunt: 20 sequential decisions at 100 milliseconds each add two seconds of delay. A model that answers in one pass with a usable probability lets apps auto-approve clear cases and escalate shaky ones, for example acting above 0.95 confidence and sending anything under 0.60 to a human.

That calibrated probability is the real product. A score of 0.9 should be right about nine times out of ten across representative cases, which is what makes thresholds meaningful.

What to do as a reader

If you run classifiers or routers, benchmark the Microsoft-Decision-1 model on your own traffic before switching. Build an eval set with rare cases and your real label mix, measure calibration with reliability plots or Brier scores, and set thresholds by the cost of false positives and negatives.

Remember the limits: it won't write, plan or explain its reasoning, and every task must reduce to fixed options. Pin versions where possible, since weights update continually. Nothing here is investment advice.

Not financial advice. DYOR, ser.

More in the pool

all ai
gm ser

Get confirmed before the crowd

Daily block at 07:00 UTC. No spam, just the block, ser.