RawLine / API docs
API reference
OpenAI-compatible. Base URL https://app.rawline.dev,
inference under /v1. Machine-readable twin:
openapi.json. Auth: Authorization: Bearer rlk_…
on every /v1/* call. Error shape:
{"error":{"type":"<code>","message":"…"}} for inference,
{"ok":false,"error":{"code":"<code>"}} for account calls.
Quickstart
cURL, plain
curl https://app.rawline.dev/v1/chat/completions \
-H "Authorization: Bearer rlk_YOUR_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-4o-mini",
"messages": [{"role": "user", "content": "Hello!"}],
"max_tokens": 256,
"temperature": 0.7
}'
Python (OpenAI SDK)
from openai import OpenAI
client = OpenAI(base_url="https://app.rawline.dev/v1", api_key="rlk_YOUR_KEY")
r = client.chat.completions.create(model="gpt-4o-mini",
messages=[{"role": "user", "content": "Hello!"}])
print(r.choices[0].message.content)
Node (OpenAI SDK)
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://app.rawline.dev/v1", apiKey: "rlk_YOUR_KEY" });
const r = await client.chat.completions.create({ model: "gpt-4o-mini",
messages: [{ role: "user", content: "Hello!" }] });
console.log(r.choices[0].message.content);
Models & prices
curl https://app.rawline.dev/v1/models -H "Authorization: Bearer rlk_YOUR_KEY"
Each entry carries id, name,
context_length (+ context_source:
members = guaranteed minimum across providers,
declared = operator plan),
max_output_tokens, supported_parameters,
architecture.input_modalities and tenant
pricing {prompt, completion} in USD per token.
Only servable models are listed: a model whose providers all lack credentials is omitted.
An unpriced model answers 402 price_not_set — price it in the
dashboard before calling.
Streaming (SSE)
curl -N https://app.rawline.dev/v1/chat/completions \
-H "Authorization: Bearer rlk_YOUR_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"gpt-4o-mini","messages":[{"role":"user","content":"Hi"}],
"stream":true, "max_tokens": 64}'
Frames are data: {…} OpenAI chunks, terminated by
data: [DONE]. Python:
stream = client.chat.completions.create(model="gpt-4o-mini",
messages=[{"role": "user", "content": "Hi"}], stream=True)
for chunk in stream:
print(chunk.choices[0].delta.content or "", end="")
Other dialects
| Route | Shape |
|---|---|
POST /v1/messages | Anthropic Messages |
POST /v1/responses | OpenAI Responses |
POST /v1beta/models/*… | Gemini :generateContent / :streamGenerateContent |
All decode into one IR, route through one failover chain and bill once.
Optional headers: X-Relay-Preset: <name> applies a preset,
X-Rawline-Session: <id> scopes conversation memory (namespaced per tenant).
Account
| Call | Notes |
|---|---|
POST /api/auth/telegram | Login Widget JSON. First login mints an API key (api_key in the reply); later logins return none. |
GET /v1/balance | balance_usd / hold_usd / spendable_usd + raw quota. |
GET+POST /v1/keys, DELETE /v1/keys/{id} | Self-service keys, max 5 live. Secret shown once. |
POST /v1/topup {"amount_usd": 25} | Heleket invoice → pay_url. Webhook credits the face amount. |
GET /v1/invoices | Own invoices, newest first, pending included. |
GET /v1/usage?limit=50 | Own spend, newest first, plus 30-day summary. Revoked keys keep their rows. |
POST /v1/promo/redeem {"code":"…"} | One redeem per user per code. |
All account calls live on the app host: cabinet.
Errors
Inference (error.type)
| HTTP | type | Meaning |
|---|---|---|
| 400 | invalid_request | Body is not JSON or the dialect could not be decoded. |
| 400 | preset_conflict | Bundle with an attached preset + X-Relay-Preset together. |
| 401 | authentication_error | Missing, unknown, revoked or expired key — one body for all. |
| 402 | insufficient_balance | Hold exceeds spendable. Top up. |
| 402 | price_not_set | Model has no operator price. Not served free. |
| 403 | model_not_allowed | Key is scoped to other models. |
| 404 | no_such_preset | Unknown X-Relay-Preset name. |
| 404 | model_not_available | No configured provider serves this model now. |
| 422 | preset_unusable | Preset could not be applied to this request. |
| 429 | rate_limit_error | Key budget spent (Retry-After set) or the provider throttles the relay. |
| 502 | provider_credential_error | Upstream refused the relay's key — operator problem, reported as 502 on purpose. |
| 502 | bad_gateway / api_error | Upstream call failed. |
| 502 | relay_encoding_error | Relay built a request the provider refused. |
| 503 | no_provider_available / overloaded_error | Nothing could serve the request. |
| 503 | billing_unavailable | Store down — refused rather than served unbilled. |
| 500 | relay_configuration_error / cache_error | Operator-side fault, detail in the relay log. |
Account (error.code)
| HTTP | code | Where |
|---|---|---|
| 401 | login_disabled / bad_signature / bad_id / stale / future / missing_hash | Telegram login |
| 404 | no_billing_account | Operator key asking for a wallet |
| 409 | key_limit | 5 live keys already |
| 404 | no_such_key | Revoke of someone else's (or nobody's) key |
| 422 | bad_amount | Top-up outside the allowed range |
| 503 | payments_disabled / payments_not_receivable | Heleket not configured / no callback URL |
| 502 | heleket_error / invoice_not_created | Gateway refused the invoice |
| 422 | invalid_code / code_expired / code_exhausted / already_redeemed | Promo redeem |
Conversation memory
Send X-Rawline-Session: <id> (any string up to 128 chars,
default default) and the relay puts the recorded assistant turns
back between the user messages you send: resend your user turns (or the whole history)
and skip the assistant ones — the relay restores them, so the provider bills fewer
repeated tokens. A request carrying only a brand-new turn matches no recorded prefix
and goes out as-is. Namespaced per tenant, so sessions never leak across wallets.
Only complete responses are stored; a truncated reply does not poison later turns.
Limits
Per-key RPM / TPM / concurrency budgets answer 429 with
Retry-After. Cache hits and single-flight joins cost no upstream
tokens and need no hold. One process per database: rate limits are per process.