Using api.saina.run
How to go from a purchase to a working call, what you are charged for, and how to retry safely. The request and response format of /v1/ask is the same as a self-hosted server; this page covers what the hosted service adds.
In short: sign up, open the sign-in link, buy credits in the console, create a key in the console, and send Authorization: Bearer $SAINA_API_KEY. You pay for input tokens; output is free. Send a UUIDv7 Idempotency-Key if you retry, and a retry is never charged twice.
- From sign-up to first call
- Authentication
- How input tokens are counted
- Credits and pricing
- The playground
- Limits
- Retries and idempotency
- Response headers
- Error codes
- Balance and usage endpoints
- What the model supports
- Data retention
- Support
From sign-up to first call
- Sign up at saina.run/signup with your email. We send a sign-in link; there is no password. A new account starts with no credits; there is no free trial.
- Open the sign-in email. Click Sign in on the page it opens. The link works once and expires after 15 minutes; if it expired, request a new one on the login page. In the console, buy credits with Stripe Checkout.
- Create a key under Keys in the console. The full key is shown once. Copy it into your secret store or an environment variable.
- Call the API:
export SAINA_API_KEY=sk_saina_... # from the console
curl https://api.saina.run/v1/ask \
-H "Authorization: Bearer $SAINA_API_KEY" \
-H 'Content-Type: application/json' \
-d '{
"model": "saina-helm-2-0.8b",
"state": "I was charged twice for my subscription this month.",
"questions": {
"team": {"type": "single_choice", "question": "Which team should handle this?",
"options": {"billing": "Billing", "technical": "Technical support"}},
"urgent": {"type": "yes_no", "question": "Does this need a reply today?"}
}
}'
The response holds the answers and usage.input_tokens. The X-Saina-Credits-Charged and X-Saina-Balance headers show what this call cost and the balance after it. The same call from Python, JavaScript or n8n is one click away in the console playground's Code panel.
Already a customer? Buy more from the Billing tab of the console; credits go to the account you are signed in to.
Authentication
- API keys look like
sk_saina_followed by 43 characters. Send one asAuthorization: Bearer …. A key can call/v1/askand/v1/systemoneand read balance, usage and request history. It cannot manage keys, buy credits or delete the account. - Shown once. We store only a hash. The console lists each key by name with the first and last four characters. If you lose a key, revoke it and create another.
- Rotation. Rotating a key issues a new one and keeps the old one working for an overlap you choose (24 hours by default, up to 7 days), so you can update every tool before it stops.
- Revocation takes effect on the next request. A revoked key gets
401 key_revoked. - The console signs you in with an emailed link instead of a password. Creating, rotating or revoking a key and deleting the account need a sign-in from the last 10 minutes; otherwise the console emails you a confirmation link first.
- Lost access? Enter your email on the login page. If an account exists, you get a sign-in link; the page answers the same way either way. From the console you can create a new key or rotate the old one.
How input tokens are counted
You pay for the input tokens in the request you send: the state once, plus each question and its options or levels, including their keys and descriptions. Text is counted as sent; objects and arrays are counted as compact JSON in the order sent, using the tokenizer published for the current price version. Prompt formatting the server adds is not charged.
The context counts once. Helm scores questions in groups of up to 8 that share one read of the state. Adding a question costs only that question and its options, however long the context is. A request with more than 8 questions runs as several groups; the context is still charged once. Each group (context, questions and options, as the model reads them) must fit in 8,192 tokens, or the request is rejected without a charge.
An illustrative example, not a measurement: a 400-token support ticket with three short questions of about 30 tokens each comes to about 400 + 3 × 30 = 490 input tokens, or about $0.00001. The console shows single requests in tokens and dollars.
Output tokens are not charged. The exact count for a request is in usage.input_tokens and in your request history. The console playground counts a request and shows the exact cost before you run it.
Credits and pricing
- Credits pay for input tokens. Usage is metered in tokens and debited from your credits, so a single request usually uses a fraction of a credit. Output tokens are free.
- Plans. Every account has one; the plan sets the credits included and the price of top-ups.
The free plan is limited to 10 requests a minute and 32 questions a request; paid plans get the API limits below. Top-ups come in $5, $10, $25, $50 and $100 packs, or any whole-dollar amount from $5 to $1,000, all at your plan's rate. Current plans and packs are also at
Plan Price Included Top-ups Free $0 25 credits a day $5 for 10,000 credits Starter $5/month 25,000 credits a month $5 for 25,000 credits Pro $10/month 55,000 credits a month $5 for 27,500 credits Max $100/month 625,000 credits a month $5 for 31,250 credits GET /v1/pricing. - Which credits go first. Credits that expire soonest: the free plan's daily credits (they expire at midnight UTC), then your plan's monthly credits (they expire when the month renews), then purchased top-ups, which never expire.
- Subscriptions renew monthly. Cancel any time in the console and keep the plan until the month ends. Upgrades start immediately and are prorated; downgrades start at renewal.
- Automatic top-up is optional and off by default: below a threshold you choose, the console buys a top-up with your saved card, up to a monthly limit you set.
- At zero balance, requests return
402 insufficient_creditsuntil credits are added. - What is charged: a completed, valid response, including a decision-mode answer that falls back (
below_threshold,below_margin,tie). A request with several questions is one operation: it finishes completely or it is not charged. - What is not charged: validation errors, rejected requests, rate limits, overload, inference failures, and replays of a result you already paid for.
- Reserved credits. Before a request runs, its counted cost is reserved, and the reservation becomes the charge when it settles. The console shows balance, reserved and available separately.
- Price versions. Each request records the price version it was charged at, so a future price change never alters past charges.
X-Saina-Price-Versionreports it. - Amounts are strings, in tokens. Balance and charge amounts in JSON (
balance,credits_chargedand the like) are decimal strings holding a 64-bit integer, for example"250000000". Parse them as integers (BigIntin JavaScript), not floats. - Invoices and receipts are in the Stripe customer portal, linked from the console's Billing tab.
The playground
The playground is in the console, after you create a free account. It sends the same /v1/ask request format to the hosted model, signed in with your console session instead of an API key.
- Runs are paid. A run is charged like an API call, at the same price, from the same credits. There is no free allowance.
- The cost is shown first. When the request is complete, the console counts it and shows the input tokens and the exact cost in dollars. Counting is free and runs no inference. The quote is tied to that exact request and expires after a short time; change the request or wait too long and it is counted again before Run is enabled.
- Charged once. Each run gets its own idempotency key. A double click or a lost connection resends the same run, which the API recognizes and does not charge again. After a run the console shows the charge and the balance at settlement, then refreshes the live balance separately.
- Visible in usage. Playground runs appear in Usage under the
playgroundchannel, with no key and the signed-in user instead. - Drafts stay in your browser. The playground keeps your drafts in this browser's storage. Inputs are never saved on our side.
- The playground has tighter limits than the API (see below). The Code panel turns the current request into curl, Python, JavaScript or an n8n node that targets
https://api.saina.runwith a$SAINA_API_KEYplaceholder.
Limits
Defaults per account. Playground traffic also counts against the API limits.
| Limit | API | Playground |
|---|---|---|
| Request body | 2 MiB | 256 KiB |
| Questions per request | 256 | 32 |
| Input tokens per request | 2,000,000 | 200,000 |
| Requests per minute | 600 | 60 (30 per session) |
| Input tokens per minute | 20,000,000 | 2,000,000 |
| Concurrent requests | 8 | 2 |
Each question's prompt must also fit the model's context window; longer inputs are rejected, never truncated. Above a limit you get 429 rate_limited with Retry-After. Cloudflare also limits each IP address to 60 requests per 10 seconds. For more capacity, talk to us.
Retries and idempotency
Without an Idempotency-Key header, every request is a new operation, and a retry after a lost response can be charged twice. To retry safely, send a key:
KEY=$(python3 -c 'import uuid; print(uuid.uuid7())') # Python 3.14+, or any UUIDv7 library curl https://api.saina.run/v1/ask \ -H "Authorization: Bearer $SAINA_API_KEY" \ -H "Idempotency-Key: $KEY" \ -H 'Content-Type: application/json' \ -d @request.json # On a timeout or a dropped connection: send the same command again, with the same $KEY and the same body.
- The key must be a UUIDv7 (RFC 9562), with its timestamp no more than 5 minutes in the future and no more than 24 hours in the past. Anything else is
invalid_request. Generate one key per logical operation, just before you first send it. - Same key, same body, already finished: you get the stored response again, with
X-Saina-Replayed: trueand no new charge. - Same key, still running:
409 request_in_progresswithRetry-After. Wait and send it again. - Same key, different body:
409 idempotency_conflict. Keys are bound to the endpoint and the body (key order and whitespace do not matter). - Same key after a failure returns the recorded failure. To try again, use a new key.
- After 24 hours the stored response is deleted. The key then returns
410 idempotency_result_expiredwith the original request ID and whether it was charged. Nothing runs again. The same 410 can appear earlier if the stored result was lost in a service incident. - After a service recovery, a key we cannot verify returns
409 idempotency_unverifiable. Start a new operation with a new key. If you generate keys ahead of time, expect this after an incident.
When to retry automatically
Retry with the same key and body only for a lost connection or unclear response, accounting_unavailable, request_in_progress, rate_limited, and overloaded or inference_unavailable when the error says "admitted": false. Honor Retry-After, back off exponentially, and stop after a few attempts. Never create a new key for a retry. Everything else needs a decision from you or your code: a new key runs the request again and may be charged again.
Response headers
| Header | Meaning |
|---|---|
X-Request-Id | On every response. Include it when you contact support. |
X-Saina-Credits-Charged | Credits charged for this call, as an integer string. |
X-Saina-Balance | The balance right after this call settled. On a replay, the balance at the original settlement. |
X-Saina-Price-Version | The price version applied. |
X-Saina-Replayed | true when this is a stored result returned for a repeated idempotency key. |
Error codes
Errors have one shape:
{"error": {"code": "insufficient_credits", "message": "...", "request_id": "req_...",
"admitted": false, "state": "not_admitted", "retry": "no"}}
state says what happened: not_admitted (nothing was recorded or reserved), in_progress, terminal (finished and will not run again under this key) or unknown (retry with the same key and body). retry is same_operation, new_operation or no.
| Code | HTTP | What to do |
|---|---|---|
invalid_api_key | 401 | Check the key. It may have a typo or belong to another environment. |
key_revoked | 401 | The key was revoked. Create a new one in the console. |
account_suspended | 403 | See suspension.next_step: buy_credits means the balance went below zero (often after a refund or dispute) and buying credits lifts it; contact_support means a security review, so email us. |
insufficient_credits | 402 | Buy credits. Nothing was charged. |
invalid_request | 400 / 422 | Fix the request; the message names the field, never its value. |
rate_limited | 429 | Wait for Retry-After, then retry the same operation. |
overloaded | 429 / 503 | With admitted: false (429), retry the same operation after Retry-After. With admitted: true (503), it was stopped and not charged; a new key is needed to run it again. |
inference_failed | 502 | The model failed. Not charged. Use a new key to try again. |
inference_unavailable | 503 | No model capacity right now. Not admitted, not charged. Retry the same operation after Retry-After. |
accounting_unavailable | 503 | Billing is briefly unavailable, so no paid work starts. Retry the same operation. |
idempotency_conflict | 409 | The key was used with a different body. Use a new key for the new request. |
request_in_progress | 409 | Still running. Retry the same operation after Retry-After. |
idempotency_result_expired | 410 | The stored result is gone. The response says whether the original was charged. It does not run again; use a new key if you need a new answer. |
idempotency_unverifiable | 409 | After a service recovery this key cannot be checked. Use a new key. |
quote_stale | 409 | Console playground only: the cost quote expired or no longer matches the request, model revision or price. Nothing was charged. Count again and run with the new quote. |
fresh_auth_required | 403 | Console only: confirm by email link, then repeat the action. |
permission_denied | 403 | The credential cannot do this; for example, an API key on a console-only route. |
not_found | 404 | Check the path or ID. |
Balance and usage endpoints
These accept an API key or a console session:
| Route | Returns |
|---|---|
GET /v1/account/balance | balance, reserved, available, overdraft_allowance, debt and any suspensions. |
GET /v1/account/usage?from=&to=&channel= | Daily totals, computed from the same records that are charged. |
GET /v1/account/requests?from=&to=&key_id=&channel=&status=&cursor=&limit= | Request history: request ID, key, channel, status, input tokens and credits for each request. |
GET /v1/pricing | Public. The current price version and packs. |
The console adds a CSV export of the request history with the same filters. channel is api, playground or openrouter, so playground spend can be told apart from API spend.
What the model supports
- Model:
saina-helm-2-0.8b(Saina Helm 2 0.8B, open weights), with a 262,144-token context, pinned to the revision thatGET /healthzreports. Requests that namesaina-helm,saina-helm-0.8borhelm-0.8bare also answered by Helm 2, and every response names the model that answered. - Question types:
yes_no,single_choice(2 to 255 options),multi_choice(multi-label, 2 to 255 options) andrating(2 to 10 levels). - Distribution mode returns a score for every option. Decision mode adds a confidence
thresholdandmin_margin, per request or per question, and areasonon every answer:accepted,below_threshold,below_marginortie. POST /v1/systemoneaccepts the System One wire format and is metered and charged the same way. Multi-label questions need/v1/ask.- The full request and response contract is in the SDK repository, and examples are on the hosted API page.
Data retention
What is kept, for how long, and where requests are processed is set out in the data policy on the hosted API page.
Support
Email [email protected]. Include the X-Request-Id of the request in question, never your API key. If a key leaked, revoke it in the console first, then tell us.