Blog · October 9, 2026

LLM Decision Making: When to Use a Chat Model and When to Use a Decision Model

LLM decision making compared: when a chat model is the right tool and when a decision model that scores your options fits, on confidence, cost and control.

Saina · October 9, 2026 · 5 min read

A lot of what happens inside automations is not writing. It's choosing. Which team handles this. Is this urgent. Which tags apply. Does this pass review. Teams reach for a chat LLM out of habit, because it's the model they already have a key for. That works, with friction. This guide is about LLM decision making: when the chat model is the right tool, and when to use a tool built for exactly this job, the decision model.

Below: what each tool is, how they differ where it matters for workflows, and a decision guide by task type. The short definition, if you don't need the full page: a decision model reads a situation and your typed questions, scores every option you supply, and returns probabilities. It never generates text.

What each tool actually is

A chat LLM generates text from a prompt. It can classify, summarize, reason and write, one completion at a time. When you use it to pick from a list, you ask in prose and parse the answer out of prose.

A decision model scores options. You pass the situation, the questions, and each question's options; you get back a probability for every option, in structured form. There is no text to parse and no way for the answer to fall outside your list, because your list is an input, not a convention.

The category is young but real on both sides of the hosting question:

Facts dated 2026-10-07; check current specs before you quote them.

The comparison that matters for workflows

Dimension Chat LLM (prompted to pick) Decision model
Output Text you parse into a category A probability for every option, keyed by your IDs
Predictability Depends on prompt, sampling, provider version Same scores for the same input and pinned weights (self-hosted); provider-managed versions (hosted)
Confidence signal Weak to none; written-out confidence isn't a probability Per-option probabilities; thresholds and margins
Calls per decision One completion per question, by convention Many typed questions in one request
Cost model Per token, prompt plus generated Usage-based (hosted) or fixed self-hosting cost
Data location Provider Provider (hosted decision model) or your servers (open weights)
Setup effort None beyond a prompt A second API (hosted) or a server to run (self-hosted)
Can it write text? Yes, that's the point No. If the step needs words written, it's the wrong tool

Two omissions from that table: accuracy and speed. Accuracy depends on the task and the models in question, so test on your own labeled set. Speed depends on deployment: Helm on CPU takes 0.9–2.7 s per request, and a hosted API on GPUs can be faster. Measure on your own setup.

A decision guide by task type

Use a chat LLM when:

Use a decision model when:

The mixed pattern most teams land on: generation steps stay on a chat model; decision steps move to a decision model; low-confidence decisions route to a human or to the chat model. That last branch is the whole point of having probabilities.

What LLM decision making looks like in code

Conceptually, one swap: the prompt-with-instructions becomes data. Several decisions about the same input go in one call, each with its own type:

from saina import Saina, SainaError

client = Saina('https://helm.internal.example', api_key='YOUR_KEY')
result = client.ask(
    model='saina-helm-0.8b',
    state='Our whole team has been locked out since the password reset email never arrived.',
    mode='decision',
    threshold=0.8,
    questions={
        'team': {
            'type': 'single_choice',
            'question': 'Which team should handle this?',
            'options': {
                'billing': 'Charges, invoices, refunds',
                'technical': 'Outages, errors, account access',
                'other': 'Anything else',
            },
        },
        'tags': {
            'type': 'multi_choice',
            'question': 'Which topics apply?',
            'threshold': 0.5,  # applied per tag; an example to tune
            'options': {
                'access': 'Login or permission trouble',
                'email': 'Emails not arriving',
                'refund': 'Money back requests',
            },
        },
        'priority': {
            'type': 'rating',
            'question': 'How urgent is this?',
            'levels': ['low', 'normal', 'high'],
        },
    },
)

team = result['answers']['team']['selection']       # e.g. 'technical', or None
tags = result['answers']['tags']['selections']      # e.g. ['access', 'email']
level = result['answers']['priority']['level']      # index into levels, or None

The single-choice answer also carries probabilities and reason; the multi-choice answer carries an independent membership score per tag; the rating returns probabilities keyed "0", "1", "2" and a zero-based level. Values in comments are illustrative.

The hosted equivalents are the same idea with different field names; the contract (probabilities per option) is what the category shares.

Handing off low-confidence decisions

In decision mode, every answer carries a reason code. accepted means the top option cleared the threshold (and the minimum margin, if you set one). below_threshold, below_margin and tie mean no selection was made, and the item should go somewhere else: a review queue, a person, or a larger chat model. Inference errors are separate: the Python SDK raises SainaError, so catch it and route the item the same way.

In n8n, the community node does this routing for you. Accepted answers leave through the Selected output; rejected ones through Fallback, with saina.reason set so the next node knows why. Inference errors stop the workflow by default; set the node's error handling to use Fallback if you want them on that branch too. Watch the Fallback rate over time: a sudden rise is the earliest sign that your inputs or options have drifted.

Where to go deeper

Checklist before you commit