Hosted API · Docs

Using api.saina.run

How to go from a purchase to a working call, what you are charged for, and how to retry safely. The request and response format of /v1/ask is the same as a self-hosted server; this page covers what the hosted service adds.

Saina · Updated October 8, 2026

In short: sign up, open the sign-in link, buy credits in the console, create a key in the console, and send Authorization: Bearer $SAINA_API_KEY. You pay for input tokens; output is free. Send a UUIDv7 Idempotency-Key if you retry, and a retry is never charged twice.

  1. From sign-up to first call
  2. Authentication
  3. How input tokens are counted
  4. Credits and pricing
  5. The playground
  6. Limits
  7. Retries and idempotency
  8. Response headers
  9. Error codes
  10. Balance and usage endpoints
  11. What the model supports
  12. Data retention
  13. Support

From sign-up to first call

  1. Sign up at saina.run/signup with your email. We send a sign-in link; there is no password. A new account starts with no credits; there is no free trial.
  2. Open the sign-in email. Click Sign in on the page it opens. The link works once and expires after 15 minutes; if it expired, request a new one on the login page. In the console, buy credits with Stripe Checkout.
  3. Create a key under Keys in the console. The full key is shown once. Copy it into your secret store or an environment variable.
  4. Call the API:
export SAINA_API_KEY=sk_saina_...   # from the console

curl https://api.saina.run/v1/ask \
  -H "Authorization: Bearer $SAINA_API_KEY" \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "saina-helm-2-0.8b",
    "state": "I was charged twice for my subscription this month.",
    "questions": {
      "team":   {"type": "single_choice", "question": "Which team should handle this?",
                 "options": {"billing": "Billing", "technical": "Technical support"}},
      "urgent": {"type": "yes_no", "question": "Does this need a reply today?"}
    }
  }'

The response holds the answers and usage.input_tokens. The X-Saina-Credits-Charged and X-Saina-Balance headers show what this call cost and the balance after it. The same call from Python, JavaScript or n8n is one click away in the console playground's Code panel.

Already a customer? Buy more from the Billing tab of the console; credits go to the account you are signed in to.

Authentication

How input tokens are counted

You pay for the input tokens in the request you send: the state once, plus each question and its options or levels, including their keys and descriptions. Text is counted as sent; objects and arrays are counted as compact JSON in the order sent, using the tokenizer published for the current price version. Prompt formatting the server adds is not charged.

The context counts once. Helm scores questions in groups of up to 8 that share one read of the state. Adding a question costs only that question and its options, however long the context is. A request with more than 8 questions runs as several groups; the context is still charged once. Each group (context, questions and options, as the model reads them) must fit in 8,192 tokens, or the request is rejected without a charge.

An illustrative example, not a measurement: a 400-token support ticket with three short questions of about 30 tokens each comes to about 400 + 3 × 30 = 490 input tokens, or about $0.00001. The console shows single requests in tokens and dollars.

Output tokens are not charged. The exact count for a request is in usage.input_tokens and in your request history. The console playground counts a request and shows the exact cost before you run it.

Credits and pricing

The playground

The playground is in the console, after you create a free account. It sends the same /v1/ask request format to the hosted model, signed in with your console session instead of an API key.

Limits

Defaults per account. Playground traffic also counts against the API limits.

LimitAPIPlayground
Request body2 MiB256 KiB
Questions per request25632
Input tokens per request2,000,000200,000
Requests per minute60060 (30 per session)
Input tokens per minute20,000,0002,000,000
Concurrent requests82

Each question's prompt must also fit the model's context window; longer inputs are rejected, never truncated. Above a limit you get 429 rate_limited with Retry-After. Cloudflare also limits each IP address to 60 requests per 10 seconds. For more capacity, talk to us.

Retries and idempotency

Without an Idempotency-Key header, every request is a new operation, and a retry after a lost response can be charged twice. To retry safely, send a key:

KEY=$(python3 -c 'import uuid; print(uuid.uuid7())')   # Python 3.14+, or any UUIDv7 library

curl https://api.saina.run/v1/ask \
  -H "Authorization: Bearer $SAINA_API_KEY" \
  -H "Idempotency-Key: $KEY" \
  -H 'Content-Type: application/json' \
  -d @request.json
# On a timeout or a dropped connection: send the same command again, with the same $KEY and the same body.

When to retry automatically

Retry with the same key and body only for a lost connection or unclear response, accounting_unavailable, request_in_progress, rate_limited, and overloaded or inference_unavailable when the error says "admitted": false. Honor Retry-After, back off exponentially, and stop after a few attempts. Never create a new key for a retry. Everything else needs a decision from you or your code: a new key runs the request again and may be charged again.

Response headers

HeaderMeaning
X-Request-IdOn every response. Include it when you contact support.
X-Saina-Credits-ChargedCredits charged for this call, as an integer string.
X-Saina-BalanceThe balance right after this call settled. On a replay, the balance at the original settlement.
X-Saina-Price-VersionThe price version applied.
X-Saina-Replayedtrue when this is a stored result returned for a repeated idempotency key.

Error codes

Errors have one shape:

{"error": {"code": "insufficient_credits", "message": "...", "request_id": "req_...",
           "admitted": false, "state": "not_admitted", "retry": "no"}}

state says what happened: not_admitted (nothing was recorded or reserved), in_progress, terminal (finished and will not run again under this key) or unknown (retry with the same key and body). retry is same_operation, new_operation or no.

CodeHTTPWhat to do
invalid_api_key401Check the key. It may have a typo or belong to another environment.
key_revoked401The key was revoked. Create a new one in the console.
account_suspended403See suspension.next_step: buy_credits means the balance went below zero (often after a refund or dispute) and buying credits lifts it; contact_support means a security review, so email us.
insufficient_credits402Buy credits. Nothing was charged.
invalid_request400 / 422Fix the request; the message names the field, never its value.
rate_limited429Wait for Retry-After, then retry the same operation.
overloaded429 / 503With admitted: false (429), retry the same operation after Retry-After. With admitted: true (503), it was stopped and not charged; a new key is needed to run it again.
inference_failed502The model failed. Not charged. Use a new key to try again.
inference_unavailable503No model capacity right now. Not admitted, not charged. Retry the same operation after Retry-After.
accounting_unavailable503Billing is briefly unavailable, so no paid work starts. Retry the same operation.
idempotency_conflict409The key was used with a different body. Use a new key for the new request.
request_in_progress409Still running. Retry the same operation after Retry-After.
idempotency_result_expired410The stored result is gone. The response says whether the original was charged. It does not run again; use a new key if you need a new answer.
idempotency_unverifiable409After a service recovery this key cannot be checked. Use a new key.
quote_stale409Console playground only: the cost quote expired or no longer matches the request, model revision or price. Nothing was charged. Count again and run with the new quote.
fresh_auth_required403Console only: confirm by email link, then repeat the action.
permission_denied403The credential cannot do this; for example, an API key on a console-only route.
not_found404Check the path or ID.

Balance and usage endpoints

These accept an API key or a console session:

RouteReturns
GET /v1/account/balancebalance, reserved, available, overdraft_allowance, debt and any suspensions.
GET /v1/account/usage?from=&to=&channel=Daily totals, computed from the same records that are charged.
GET /v1/account/requests?from=&to=&key_id=&channel=&status=&cursor=&limit=Request history: request ID, key, channel, status, input tokens and credits for each request.
GET /v1/pricingPublic. The current price version and packs.

The console adds a CSV export of the request history with the same filters. channel is api, playground or openrouter, so playground spend can be told apart from API spend.

What the model supports

Data retention

What is kept, for how long, and where requests are processed is set out in the data policy on the hosted API page.

Support

Email [email protected]. Include the X-Request-Id of the request in question, never your API key. If a key leaked, revoke it in the console first, then tell us.