api.saina.run
Saina Helm, served by us, for teams that want to try it or run it from tools that can't host a model, such as n8n Cloud, Make and Zapier. Same contract as a self-hosted server. Request and response bodies are never stored.
In short: https://api.saina.run runs the same open server you can self-host, on our hardware, with a zero-retention policy. Access is by API key: create a free account, buy credits and create one in the console, or self-host with no key at all.
Endpoints
| Method and path | What it does |
|---|---|
GET /healthz | Readiness, the model name and the pinned weights revision, so you can verify which weights answered you. No key needed. |
POST /v1/ask | Saina's native contract: typed questions (yes_no, single_choice, multi_choice, rating), decision mode with a confidence threshold and minimum margin, and a reason code on every answer. Documented in the SDK. |
POST /v1/systemone | The System One wire format used by Jev clients and OpenRouter's Decisions API (noul, choice, score). Point an existing System One client at this base URL and it works. Multi-label questions aren't part of that format; use /v1/ask for those. |
Both POST routes take a bearer token: Authorization: Bearer $SAINA_API_KEY. Each call is charged in credits for its input tokens. Errors have a stable code and say whether a retry is safe; send an Idempotency-Key to retry without being charged twice. Details: API docs.
Examples
Native contract
curl https://api.saina.run/v1/ask \
-H "Authorization: Bearer $SAINA_API_KEY" -H 'Content-Type: application/json' \
-d '{
"model": "saina-helm-2-0.8b",
"mode": "decision", "threshold": 0.8, "min_margin": 0.05,
"state": "I was charged twice for last month and nobody answers the chat.",
"questions": {
"team": {"type": "single_choice", "question": "Which team should handle this?",
"options": {"billing": "Payments and refunds", "technical": "Bugs and outages", "other": null}},
"tags": {"type": "multi_choice", "question": "Which tags apply?",
"options": {"refund": null, "duplicate_charge": null, "angry": null, "outage": null}},
"priority": {"type": "rating", "question": "How urgent is this?", "levels": ["low", "normal", "high"]}
}
}'
Every answer carries the probability of each option and a reason: accepted, below_threshold, below_margin or tie. An accepted answer clears your configured thresholds; it is not a guarantee of correctness. Apply your own checks and any required human review before acting, and route uncertain answers to a fallback.
System One format
curl https://api.saina.run/v1/systemone \
-H "Authorization: Bearer $SAINA_API_KEY" -H 'Content-Type: application/json' \
-d '{
"model": "saina-helm-2-0.8b",
"state": "Help! My payouts have been failing for 3 days.",
"questions": {
"is_urgent": {"type": "noul", "instructions": "Does this message convey urgency?"},
"team": {"type": "choice", "instructions": "Which team?", "criteria": {"billing": "Money", "technical": "Bugs"}},
"severity": {"type": "score", "instructions": "How severe?", "criteria": ["low", "medium", "high"]}
}
}'
Python
from saina import Saina
client = Saina('https://api.saina.run', api_key='YOUR_KEY')
result = client.ask(model='saina-helm-2-0.8b', state='I was charged twice.', questions={
'issue': {'type': 'single_choice', 'question': 'Identify the issue.',
'options': {'duplicate': 'duplicate charge', 'delivery': 'card delivery'}}})
The console playground (after you create a free account) builds these requests for you and copies them as curl, Python, JavaScript or n8n. Credits, keys, limits, retries and error codes are in the API docs.
Data policy: zero retention
This policy describes inference-content handling by the hosted API today. Website, waitlist, support and browser storage are described in section 4 of the terms. Material changes follow the notice and acceptance process; changing this page's date alone does not change an existing customer's agreement. Paid-plan retention described in the terms is not enabled by this policy.
- Request and response bodies are never stored. The server keeps nothing between requests: no database, no disk writes, no cache of your text, no analytics on content. Inference runs in memory and the result is returned.
- No per-request logs on our side. The server runs with access logging off. It logs startup and errors only, and those lines never include request content. Error responses name the field that failed, not its value.
- Not used for training. Nothing you send is used to train, evaluate or tune any model.
- Transit.
api.saina.runis reached through Cloudflare, which terminates TLS and forwards the request to our server over a private tunnel. Cloudflare sees requests in transit and keeps its own edge metadata (timestamps, status codes, request counts) under its policies; it does not store bodies for us, and we have not enabled any Cloudflare feature that logs them. If you need no intermediary at all, self-host: it's the same server. - Keys. A key identifies a caller and can be revoked individually. We don't tie keys to request content, because there is no request content kept to tie them to.
- Location. Requests are processed in Canada (British Columbia). Any standby capacity runs under the same policy, and its region is shown on this page while in use. Weights are the public Saina Helm 2 0.8B release, pinned to the revision reported by
/healthz, and are only changed deliberately. - Through OpenRouter. Requests made through OpenRouter pass through OpenRouter under its data policy before reaching this server, where the rules above apply.
Questions about any of this: [email protected]. Use of the hosted API is under the terms of service.
Limits and support
- Up to 256 questions per request. Questions are scored in groups of up to 8 that share one read of the context; each group (context, questions and options) must fit in 8,192 tokens. Longer inputs are rejected, never truncated, and are not charged.
- You pay for the input tokens you send: the context once per request, plus each question and its options. Prompt formatting and output are not charged. Details are in the terms.
- For guaranteed capacity or an on-premise deployment, talk to us.
- You can self-host the published server with the same request contract: run-saina/deploy (Docker) or
pip install 'saina[local,server]'. Hardware, model versions, capacity and operational behavior may differ.
Get a key
Create a free account: the free plan includes 25 credits a day, and output is free. For more, subscribe from $5 a month or buy top-ups in the console with Stripe. Keys are created in the console.