EN / RUSign inGet API key

RawLine / API docs

API reference

OpenAI-compatible. Base URL https://app.rawline.dev, inference under /v1. Machine-readable twin: openapi.json. Auth: Authorization: Bearer rlk_… on every /v1/* call. Error shape: {"error":{"type":"<code>","message":"…"}} for inference, {"ok":false,"error":{"code":"<code>"}} for account calls.

Quickstart

cURL, plain

curl https://app.rawline.dev/v1/chat/completions \
  -H "Authorization: Bearer rlk_YOUR_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-4o-mini",
    "messages": [{"role": "user", "content": "Hello!"}],
    "max_tokens": 256,
    "temperature": 0.7
  }'

Python (OpenAI SDK)

from openai import OpenAI
client = OpenAI(base_url="https://app.rawline.dev/v1", api_key="rlk_YOUR_KEY")
r = client.chat.completions.create(model="gpt-4o-mini",
    messages=[{"role": "user", "content": "Hello!"}])
print(r.choices[0].message.content)

Node (OpenAI SDK)

import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://app.rawline.dev/v1", apiKey: "rlk_YOUR_KEY" });
const r = await client.chat.completions.create({ model: "gpt-4o-mini",
  messages: [{ role: "user", content: "Hello!" }] });
console.log(r.choices[0].message.content);

Models & prices

curl https://app.rawline.dev/v1/models -H "Authorization: Bearer rlk_YOUR_KEY"

Each entry carries id, name, context_length (+ context_source: members = guaranteed minimum across providers, declared = operator plan), max_output_tokens, supported_parameters, architecture.input_modalities and tenant pricing {prompt, completion} in USD per token. Only servable models are listed: a model whose providers all lack credentials is omitted. An unpriced model answers 402 price_not_set — price it in the dashboard before calling.

Streaming (SSE)

curl -N https://app.rawline.dev/v1/chat/completions \
  -H "Authorization: Bearer rlk_YOUR_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"gpt-4o-mini","messages":[{"role":"user","content":"Hi"}],
       "stream":true, "max_tokens": 64}'

Frames are data: {…} OpenAI chunks, terminated by data: [DONE]. Python:

stream = client.chat.completions.create(model="gpt-4o-mini",
    messages=[{"role": "user", "content": "Hi"}], stream=True)
for chunk in stream:
    print(chunk.choices[0].delta.content or "", end="")

Other dialects

RouteShape
POST /v1/messagesAnthropic Messages
POST /v1/responsesOpenAI Responses
POST /v1beta/models/*…Gemini :generateContent / :streamGenerateContent

All decode into one IR, route through one failover chain and bill once. Optional headers: X-Relay-Preset: <name> applies a preset, X-Rawline-Session: <id> scopes conversation memory (namespaced per tenant).

Account

CallNotes
POST /api/auth/telegramLogin Widget JSON. First login mints an API key (api_key in the reply); later logins return none.
GET /v1/balancebalance_usd / hold_usd / spendable_usd + raw quota.
GET+POST /v1/keys, DELETE /v1/keys/{id}Self-service keys, max 5 live. Secret shown once.
POST /v1/topup {"amount_usd": 25}Heleket invoice → pay_url. Webhook credits the face amount.
GET /v1/invoicesOwn invoices, newest first, pending included.
GET /v1/usage?limit=50Own spend, newest first, plus 30-day summary. Revoked keys keep their rows.
POST /v1/promo/redeem {"code":"…"}One redeem per user per code.

All account calls live on the app host: cabinet.

Errors

Inference (error.type)

HTTPtypeMeaning
400invalid_requestBody is not JSON or the dialect could not be decoded.
400preset_conflictBundle with an attached preset + X-Relay-Preset together.
401authentication_errorMissing, unknown, revoked or expired key — one body for all.
402insufficient_balanceHold exceeds spendable. Top up.
402price_not_setModel has no operator price. Not served free.
403model_not_allowedKey is scoped to other models.
404no_such_presetUnknown X-Relay-Preset name.
404model_not_availableNo configured provider serves this model now.
422preset_unusablePreset could not be applied to this request.
429rate_limit_errorKey budget spent (Retry-After set) or the provider throttles the relay.
502provider_credential_errorUpstream refused the relay's key — operator problem, reported as 502 on purpose.
502bad_gateway / api_errorUpstream call failed.
502relay_encoding_errorRelay built a request the provider refused.
503no_provider_available / overloaded_errorNothing could serve the request.
503billing_unavailableStore down — refused rather than served unbilled.
500relay_configuration_error / cache_errorOperator-side fault, detail in the relay log.

Account (error.code)

HTTPcodeWhere
401login_disabled / bad_signature / bad_id / stale / future / missing_hashTelegram login
404no_billing_accountOperator key asking for a wallet
409key_limit5 live keys already
404no_such_keyRevoke of someone else's (or nobody's) key
422bad_amountTop-up outside the allowed range
503payments_disabled / payments_not_receivableHeleket not configured / no callback URL
502heleket_error / invoice_not_createdGateway refused the invoice
422invalid_code / code_expired / code_exhausted / already_redeemedPromo redeem

Conversation memory

Send X-Rawline-Session: <id> (any string up to 128 chars, default default) and the relay puts the recorded assistant turns back between the user messages you send: resend your user turns (or the whole history) and skip the assistant ones — the relay restores them, so the provider bills fewer repeated tokens. A request carrying only a brand-new turn matches no recorded prefix and goes out as-is. Namespaced per tenant, so sessions never leak across wallets. Only complete responses are stored; a truncated reply does not poison later turns.

Limits

Per-key RPM / TPM / concurrency budgets answer 429 with Retry-After. Cache hits and single-flight joins cost no upstream tokens and need no hold. One process per database: rate limits are per process.