LLM Decision Making: When to Use a Chat Model and When to Use a Decision Model
LLM decision making compared: when a chat model is the right tool and when a decision model that scores your options fits, on confidence, cost and control.
A lot of what happens inside automations is not writing. It's choosing. Which team handles this. Is this urgent. Which tags apply. Does this pass review. Teams reach for a chat LLM out of habit, because it's the model they already have a key for. That works, with friction. This guide is about LLM decision making: when the chat model is the right tool, and when to use a tool built for exactly this job, the decision model.
Below: what each tool is, how they differ where it matters for workflows, and a decision guide by task type. The short definition, if you don't need the full page: a decision model reads a situation and your typed questions, scores every option you supply, and returns probabilities. It never generates text.
What each tool actually is
A chat LLM generates text from a prompt. It can classify, summarize, reason and write, one completion at a time. When you use it to pick from a list, you ask in prose and parse the answer out of prose.
A decision model scores options. You pass the situation, the questions, and each question's options; you get back a probability for every option, in structured form. There is no text to parse and no way for the answer to fall outside your list, because your list is an input, not a convention.
The category is young but real on both sides of the hosting question:
- Hosted: GPT-6 Luna Decisions (OpenAI, released 2026-10-06): GPT-6 Luna behind a Decisions API; yes/no, choice and rubric-score questions with probabilities, up to 200 per request; state can be text, JSON or images. Its own request shape; chat SDKs don't apply. Check OpenAI's pricing page for the current rate.
- Self-hosted: Saina Helm: a 0.8B open-weights decision model; single-choice, multi-label, yes/no and rating questions, 1–256 per
/v1/askrequest; decision mode with a confidence threshold (default 0.8) and optional minimum margin, set per request or per question; weights pinned to a Hub revision; runs on your own servers via Docker or pip (~4 GB RAM on CPU; measured 0.9–2.7 s per request, local Docker test, 2026-10-06). n8n community node with Selected/Fallback outputs on self-hosted n8n; HTTP template on n8n Cloud. Also on Ollama and as MLX for Apple silicon. The same server is available hosted atapi.saina.run(free plan: 25 credits a day).
Facts dated 2026-10-07; check current specs before you quote them.
The comparison that matters for workflows
| Dimension | Chat LLM (prompted to pick) | Decision model |
|---|---|---|
| Output | Text you parse into a category | A probability for every option, keyed by your IDs |
| Predictability | Depends on prompt, sampling, provider version | Same scores for the same input and pinned weights (self-hosted); provider-managed versions (hosted) |
| Confidence signal | Weak to none; written-out confidence isn't a probability | Per-option probabilities; thresholds and margins |
| Calls per decision | One completion per question, by convention | Many typed questions in one request |
| Cost model | Per token, prompt plus generated | Usage-based (hosted) or fixed self-hosting cost |
| Data location | Provider | Provider (hosted decision model) or your servers (open weights) |
| Setup effort | None beyond a prompt | A second API (hosted) or a server to run (self-hosted) |
| Can it write text? | Yes, that's the point | No. If the step needs words written, it's the wrong tool |
Two omissions from that table: accuracy and speed. Accuracy depends on the task and the models in question, so test on your own labeled set. Speed depends on deployment: Helm on CPU takes 0.9–2.7 s per request, and a hosted API on GPUs can be faster. Measure on your own setup.
A decision guide by task type
Use a chat LLM when:
- The step produces text: replies, summaries, extractions, drafts.
- The question needs multi-step reasoning or outside knowledge ("read this contract and flag unusual terms").
- The option set is open-ended or changes per item ("name the product mentioned").
- Volume is low and a hosted cheap tier is the simplest thing that works.
Use a decision model when:
- The step picks between options you already know: route, tag, gate, rate, escalate.
- You need a confidence signal to act on: thresholds, review queues, human handoff.
- Many related decisions share one situation (team + priority + tags) and belong in one call.
- The option list changes often, or differs per tenant, and you don't want to re-engineer prompts.
The mixed pattern most teams land on: generation steps stay on a chat model; decision steps move to a decision model; low-confidence decisions route to a human or to the chat model. That last branch is the whole point of having probabilities.
What LLM decision making looks like in code
Conceptually, one swap: the prompt-with-instructions becomes data. Several decisions about the same input go in one call, each with its own type:
from saina import Saina, SainaError
client = Saina('https://helm.internal.example', api_key='YOUR_KEY')
result = client.ask(
model='saina-helm-0.8b',
state='Our whole team has been locked out since the password reset email never arrived.',
mode='decision',
threshold=0.8,
questions={
'team': {
'type': 'single_choice',
'question': 'Which team should handle this?',
'options': {
'billing': 'Charges, invoices, refunds',
'technical': 'Outages, errors, account access',
'other': 'Anything else',
},
},
'tags': {
'type': 'multi_choice',
'question': 'Which topics apply?',
'threshold': 0.5, # applied per tag; an example to tune
'options': {
'access': 'Login or permission trouble',
'email': 'Emails not arriving',
'refund': 'Money back requests',
},
},
'priority': {
'type': 'rating',
'question': 'How urgent is this?',
'levels': ['low', 'normal', 'high'],
},
},
)
team = result['answers']['team']['selection'] # e.g. 'technical', or None
tags = result['answers']['tags']['selections'] # e.g. ['access', 'email']
level = result['answers']['priority']['level'] # index into levels, or None
The single-choice answer also carries probabilities and reason; the multi-choice answer carries an independent membership score per tag; the rating returns probabilities keyed "0", "1", "2" and a zero-based level. Values in comments are illustrative.
The hosted equivalents are the same idea with different field names; the contract (probabilities per option) is what the category shares.
Handing off low-confidence decisions
In decision mode, every answer carries a reason code. accepted means the top option cleared the threshold (and the minimum margin, if you set one). below_threshold, below_margin and tie mean no selection was made, and the item should go somewhere else: a review queue, a person, or a larger chat model. Inference errors are separate: the Python SDK raises SainaError, so catch it and route the item the same way.
In n8n, the community node does this routing for you. Accepted answers leave through the Selected output; rejected ones through Fallback, with saina.reason set so the next node knows why. Inference errors stop the workflow by default; set the node's error handling to use Fallback if you want them on that branch too. Watch the Fallback rate over time: a sudden rise is the earliest sign that your inputs or options have drifted.
Where to go deeper
- The definition page for the category, kept current: What is a decision model?
Checklist before you commit
- List every LLM step in the workflow; mark each as generation or decision.
- For decision steps: write out the options. If you can't, it's a generation step.
- Choose hosted vs. self-hosted on data location and cost model, not on accuracy claims.
- Build a labeled set; run the candidate in shadow mode against current answers.
- Set a threshold and margin; decide what happens below the bar (human? chat model? queue?).
- Pin the model version; schedule test-set re-runs after any change.