Explainer

What is a decision model?

A decision model is an AI model that answers questions about a situation by scoring the options you give it, instead of writing text. This page explains what that means, how it differs from a chat LLM and from a classifier, and when to use one in a workflow.

Saina · Updated October 7, 2026 · 8 min read

In short: you give a decision model the situation, your questions and your options. It returns a probability for every option, in one pass, with no text to parse. It is built for the steps in a workflow that pick what happens next: route, prioritize, tag, check, escalate.

The definition

A decision model reads a state (any text or structured record) and answers typed questions about it. Each question comes with its own options. The model scores every option and returns the scores as structured output. It does not generate text.

Three things in that definition do the work.

What a request looks like

The situation is a support message. The question asks which team should handle it. The options are the teams.

{
  "state": "I was charged twice for my subscription this month.",
  "questions": {
    "team": {
      "type": "single_choice",
      "question": "Which team should handle this?",
      "options": {
        "billing":   "Charges, invoices, refunds",
        "technical": "Outages, errors, account access",
        "other":     "Anything else"
      }
    }
  }
}

Back comes one keyed answer, with a score for every option.

{
  "team": {
    "type": "single_choice",
    "selection": "billing",
    "probabilities": { "billing": 0.94, "technical": 0.04, "other": 0.02 }
  }
}

Nothing to parse, nothing to validate, nothing the model could have invented. If the top two scores had been close, that would be visible too, and the case could go to a person.

Decision model vs. chat LLM

A chat model can classify. It just wasn't built for it. When you ask a chat LLM to "reply with billing, technical or other," you are borrowing a text generator for a scoring job, and the mismatch shows up in production.

Chat LLMDecision model
OutputText you parse into a categoryA score for every option, keyed by your IDs
Invented answersPossible: a label you never definedNot possible: it can only score your options
UncertaintyHidden; you get a pickExplicit; you get a probability per option
Prompt driftReword the prompt or let the provider update, and answers changeNo prompt; a pinned model gives the same answer to the same input
Cost per decisionA full chat completionOne scoring pass, no text generated
LatencySecondsTens to hundreds of milliseconds
WritingYesNo. It scores; it doesn't generate

The last row is the trade. If the step needs words written, you need a language model. If the step needs a pick, a decision model is the right shape of tool.

Decision model vs. text classifier

Before LLMs, this job belonged to text classifiers: a model fine-tuned on a fixed set of labels. They are fast and cheap, and they still work. Their limit is the fixed label set.

Think of a decision model as a zero-shot classifier that also understands the question, not just the labels, and that returns a probability for every option rather than one winner.

The four question types

TypeQuestionAnswer
yes_noDoes this need a human?P(yes) and P(no)
single_choiceWhich team should handle this?One probability per option, summing to one
multi_choiceWhich tags apply?An independent score per tag; several can qualify
ratingHow severe is this, on your scale?A probability per level and the expected level

Most decision models answer yes/no, single-choice and rating questions. Multi-select is less common and matters for tagging, where more than one answer is usually right.

Scores, thresholds and the human in the loop

The score is the part that changes how you build. A keyword rule or a chat pick gives you a yes. A decision model gives you a yes and a number, and the number lets you set a bar.

  1. Set a threshold. Act only when the top option clears it.
  2. Set a minimum gap. If the top two options are too close to call, don't act.
  3. Route the rest to a person. Cases below the bar come back as an explicit fallback instead of a pick, with the scores attached.

That is what makes a decision model safe to put in front of an action. Nothing is dropped silently, and the cases that need judgement get it. Set the bar from your own data: a score is a measure of the model's view, not a guarantee of correctness.

Many decisions, one situation

Workflows rarely ask one question. A support message needs a team, a priority, a sentiment and maybe a flag for a refund request. A decision model takes all the questions in one request and answers them together from the same reading of the situation. One call replaces a chain of separate ones, and the answers are consistent with each other.

When to use one

Good fits

Not a fit

Where the term comes from

"Decision model" became a category of its own in September 2026, when the first hosted models built around this contract appeared and were quickly followed by open-weight ones from several labs. The large platforms followed within weeks: Ollama added a "decision" capability, OpenRouter added a decisions output type, and on October 6, 2026 OpenAI released GPT-6 Luna Decisions, its GPT-6 Luna model served through a dedicated Decisions API. The idea has older roots: scoring candidates with a single pass of an encoder, zero-shot classification, and reward models all do part of this job. What is new is the shared shape: a state, typed questions with their options, and probabilities back, served over a common API so clients work across models. Model registries now list "decision" alongside "vision" and "embedding" as a capability.

Where Saina fits

Decision models now come in two forms. Hosted ones, such as OpenAI's GPT-6 Luna Decisions, run on the provider's servers and bill per token, so there is nothing to operate. Open-weight ones run wherever you put them: the data never leaves your infrastructure, the cost doesn't grow with every call, and you decide when the model changes.

Saina builds open decision models for automated workflows. Helm, the first, is a 0.8B model you can self-host, call over HTTP, or drop into n8n as a node with a built-in fallback output. It answers all four question types, including multi-label tagging, up to 256 questions in one request, and returns an explicit fallback with the reason when a call is below your bar.

Frequently asked questions

What is a decision model?

An AI model that reads a situation and answers typed questions about it with a score for every option you supply, instead of writing text. You define the options at request time; it returns a probability for each one in a single pass.

How is a decision model different from a chat LLM?

A chat LLM generates text that you then parse into a category. A decision model never generates text: it reads the options you gave it and scores them directly, so the output is always one of your options, always structured, and comes with a probability you can set a threshold on.

How is a decision model different from a text classifier?

A classic classifier is trained on a fixed set of labels and has to be retrained when they change. A decision model takes the labels as part of each request, described in plain language, so one model handles routing, tagging, ratings and yes/no checks without per-task training.

When should I use a decision model?

At any step in a workflow where the job is to pick between options you already know: route a ticket, set a priority, tag feedback, check a rubric, decide whether to escalate. Not when the step needs words written.

Can a decision model say it isn't sure?

Yes. Every option comes back with a probability, so you can set a threshold and a minimum gap between the top two options. When a case falls below the bar, the model returns an explicit fallback instead of a pick, and the case can go to a person.

Do I need to train a decision model for my task?

No. The options and the question are part of each request. Describe them in plain words and the model scores them. Fine-tuning on your own data can raise accuracy further, but it isn't required to start.

Which decision models can I use?

Hosted ones from API providers, such as OpenAI's GPT-6 Luna Decisions, and open-weight ones you run yourself, such as Saina Helm. Pick hosted if you want nothing to operate; pick open-weight if the data must stay in your infrastructure or you want a fixed cost and control over model versions.

Can I run a decision model on my own servers?

Open-weight decision models, including Saina Helm, run on your own infrastructure. Helm is 0.8B parameters and runs on a CPU or a small GPU, so data never has to leave your organization.