Introducing Saina Helm
Saina is the decision layer for automation: models built only for the choosing step in a workflow. Helm is the first of them, and the best decision model under 1 billion parameters for intent classification and routing. Open weights, and v1.0 is out today.
A lot of what automations do isn't writing. It's choosing. Which team takes this ticket, whether it needs a reply today, which tags apply, whether this invoice can be paid without a second look. Most teams answer those questions with a chat model and a parser, because that's the model they already have.
Saina is the decision layer for automation: models built only for that choosing step. You describe the situation, ask typed questions, list the options, and get back a probability for every option. Nothing is generated, so there's no free text to parse and no answer that falls outside your list. If you're new to the idea, the decision model page covers the category, including hosted options.
Saina Helm is the first of them: the best decision model under 1 billion parameters, with open weights, and v1.0 is out today. Best depends on the use case, of course. Where Helm is best is intent classification and routing, where it beats models many times its size; the benchmarks show where it stands and what's still in training.
Why a family
Different decisions need different trade-offs: a small model you can run on a CPU next to your data, larger ones for harder judgment calls. We expect to ship more than one model, and we'd rather you never have to rewrite a workflow to switch between them.
So every Saina model will speak the same contract: /v1/ask, the same four question types, the same answer shape. Helm is the first model to implement it. When the next one arrives, changing the model field should be the only change you make.
What Helm does
Helm takes a state (a ticket, an email, a JSON record, anything you can put in text) and between 1 and 256 questions in a single call. Each question has one of four types:
yes_no: does this need a reply today?single_choice: which team should handle it? (2–255 options)multi_choice: which tags apply? Each option is scored independently, so several can be true at once.rating: how urgent is it, on a scale you define with 2–10 levels?
Every answer comes back with its probabilities. In decision mode, Helm also tells you whether it's sure enough to act on, which matters more than the answer itself.
Here's a real ticket that's half billing problem and half login problem:
from saina import Saina
client = Saina('https://your-saina-server', api_key='...')
result = client.ask(
model='saina-helm-0.8b',
mode='decision',
threshold=0.8,
state="Hi, I was charged twice for my October invoice and I can't log in "
"to download the receipt. Our finance close is Friday.",
questions={
'team': {
'type': 'single_choice',
'question': 'Which team should handle this ticket?',
'options': {
'billing': 'Billing: charges, refunds, invoices',
'technical': 'Technical support: login, bugs, outages',
'sales': 'Sales: plans, upgrades, quotes',
},
},
'urgent': {'type': 'yes_no', 'question': 'Does this need a reply today?'},
'tags': {
'type': 'multi_choice',
'question': 'Which topics does the ticket mention?',
'threshold': 0.5,
'options': {
'duplicate_charge': 'Duplicate charge',
'login': 'Login problem',
'receipt': 'Receipt or invoice request',
'cancellation': 'Cancellation',
},
},
},
)
And what came back (real output from the demo Space on 2026-10-07, rounded):
{
"team": { "selection": null, "reason": "below_threshold",
"probabilities": { "billing": 0.42, "technical": 0.58, "sales": 0.00 } },
"urgent": { "selected": null, "reason": "below_threshold",
"yes": 0.60, "no": 0.40 },
"tags": { "selections": ["duplicate_charge", "receipt"], "reason": "accepted",
"memberships": { "duplicate_charge": 0.97, "login": 0.38,
"receipt": 0.72, "cancellation": 0.04 } }
}
(The answer field differs by type: selection for choices, selected for yes/no, selections for tags.)
Helm leaned towards technical support but split 58/42 with billing, which is fair for this ticket. A chat model would have picked one and moved on. Helm returned no selection and a reason, so the workflow can send the ticket to a person. The urgency call is also too close to automate. The tags are confident enough to apply. They missed the login problem (0.38), which is the kind of thing you find by testing on your own tickets, not by trusting a demo.
That's the core of the design:
- A threshold and a minimum margin for each request or each question. The default threshold is 0.8.
- A reason on every answer:
accepted,below_threshold,below_marginortie. - In n8n, the Saina node sends accepted answers to its Selected output and everything else to Fallback, so the human review step is just a branch in your workflow.
Where it runs
Helm is meant to run where your data already is.
Open weights on Hugging Face:
run-saina/saina-helm-0.8b, plus an MLX build for Apple silicon.Your own server: one container, about 4 GB of RAM on CPU. Requests took 0.9–2.7 s on CPU in our local Docker test (2026-10-06).
export SAINA_API_KEY=$(openssl rand -hex 32) docker run -d --name saina -p 127.0.0.1:8000:8000 \ -e SAINA_API_KEY=$SAINA_API_KEY \ -e SAINA_CHECKPOINT=run-saina/saina-helm-0.8b \ -e SAINA_REVISION=e82055b348d02bb43552ccba88eda24835ba6baf \ -v saina-hf-cache:/cache --memory 8g --restart unless-stopped \ ghcr.io/run-saina/saina-server:0.1.2The first start downloads the pinned weights (about 1.7 GB). Prefer Python?
pip install 'saina[local,server]'runs the same server withuvicorn saina.server:create_app --factory. Compose files and cloud notes are inrun-saina/deploy.Ollama (0.40 or later):
ollama pull run-saina/saina-helm-0.8b. It serves the System One decision format. Multi-label tags aren't available there yet.Hosted:
https://api.saina.runserves the same API under a zero-retention policy, for tools that can't host a model (n8n Cloud, Make, Zapier). Create a free account to try it: the free plan includes 25 credits a day.n8n: the community node
@run-saina/n8n-nodes-sainaworks on self-hosted n8n. On n8n Cloud, use the HTTP Request template instead. Four ready-made workflows (support routing, model escalation, invoice review and the portable HTTP version) are inrun-saina/n8n-nodes.SDKs: Python (
saina) and JavaScript (@run-saina/sdk).
The n8n routing guide walks through the support-routing workflow end to end.
You decide when the model changes
Helm v1.0 is pinned: tag v1.0 on Hugging Face and 0.8b-v1.0 on Ollama. A new version only reaches you when you choose it, and the version numbers mean something:
- v1.x: retrained weights, same labels, same request format. Re-test your thresholds, but nothing in your workflow has to change.
- v2.0: anything that would make you change a request.
The same rule will apply to every Saina model.
What Helm v1.0 doesn't do yet
- Calibration hasn't been measured. Helm's probabilities rank options, but don't read 0.9 as "right 90% of the time" until we publish that measurement. Tune thresholds on a labeled sample of your own data.
- CPU is the measured path. Expect seconds per request, not milliseconds. One server process handles one request at a time.
- Multi-label tags need the
/v1/askserver. Ollama's decision format doesn't support them. - The GGUF build is experimental, and there is no ONNX release yet.
- Input is capped at 8,192 tokens by default. The cap applies to each question separately: the state, that question and its options.
What's next
- An open benchmark. We're running every sub-1B decision model we can get, through public endpoints and weights, on the same machine with the same prompts. We'll publish the harness, the data and the date, including the results where Helm loses.
- More Saina models. The next one is in training. We'll announce it when it's ready to ship, not before.
Try it
- No install: the demo on Hugging Face. Paste a situation, ask questions, see the probabilities.
- Against your own server: the playground.
- In a workflow: start from the support-routing template in
run-saina/n8n-nodes. - To weigh it against a chat model: read LLM decision making: when to use a chat model and when to use a decision model.
Questions or integration help: [email protected].