Documentation
Cognia API
Agent Keys call the same models as the chat workspace through an OpenAI-compatible API at https://usecognia.xyz/v1. OpenAI SDKs work by changing the base URL and key. The supported request fields are listed below; anything else is rejected with a clear error instead of being ignored.
Quickstart
- Sign in with email or a wallet.
- Add funds on the Billing page with USDG on Robinhood Chain.
- Create a key on the Agent Keys page. It is shown once, so store it in a secret manager.
- Pick a model id from Models or
GET /v1/models.
export COGNIA_API_KEY="sk-cog-..."
curl https://usecognia.xyz/v1/chat/completions \
-H "Authorization: Bearer $COGNIA_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "ag/claude-opus-4-6-thinking",
"max_tokens": 300,
"messages": [{"role": "user", "content": "Write a haiku about ledgers."}]
}'import os
from openai import OpenAI
client = OpenAI(base_url="https://usecognia.xyz/v1", api_key=os.environ["COGNIA_API_KEY"])
resp = client.chat.completions.create(
model="ag/claude-opus-4-6-thinking",
max_tokens=300,
messages=[{"role": "user", "content": "Write a haiku about ledgers."}],
)
print(resp.choices[0].message.content)import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://usecognia.xyz/v1", apiKey: process.env.COGNIA_API_KEY });
const resp = await client.chat.completions.create({
model: "ag/claude-opus-4-6-thinking",
max_tokens: 300,
messages: [{ role: "user", content: "Write a haiku about ledgers." }],
});
console.log(resp.choices[0].message.content);Authentication
Send your key as Authorization: Bearer sk-cog-.... The Messages endpoint also accepts x-api-key. Keys are stored only as keyed hashes, so a lost key cannot be recovered: revoke it and create a new one. Browser sessions are not accepted on /v1, and Agent Keys are not accepted by the web app's own endpoints. Never put a key in client-side code.
API reference
| Endpoint | Purpose |
|---|---|
| GET /v1/models | OpenAI-format model list. Each model has a cognia object with prices in USD per 1M tokens and capabilities. Key optional. |
| POST /v1/chat/completions | OpenAI Chat Completions. Supports stream, tools, tool_choice, response_format, stop, temperature, top_p, seed, max_tokens or max_completion_tokens. n must be 1. |
| POST /v1/messages | Anthropic Messages format for models marked "Messages API". |
| POST /v1/images/generations | Not available yet. Returns 501. |
If you omit max_tokens, the model's default output limit applies, lowered if needed to what your balance and the key's budget can cover; the limit actually used is returned in x-cognia-max-tokens-applied. A max_tokens you set yourself is never changed: if it cannot be covered, the request is refused with 402. Output is always bounded so the amount held is a true maximum.
Streaming
Set "stream": true to receive Server-Sent Events. Pass stream_options: {"include_usage": true} to get the final usage chunk, exactly as with OpenAI.
stream = client.chat.completions.create(
model="ag/claude-opus-4-6-thinking",
stream=True,
stream_options={"include_usage": True},
messages=[{"role": "user", "content": "Count to five."}],
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="")Tool calls
Tool definitions and tool call deltas pass through unchanged for models that support tools (see the Tools badge). Your code executes the tool and sends the result back as a tool message.
const resp = await client.chat.completions.create({
model: "ag/claude-opus-4-6-thinking",
messages: [{ role: "user", content: "What's the weather in Jakarta?" }],
tools: [{
type: "function",
function: {
name: "get_weather",
parameters: { type: "object", properties: { city: { type: "string" } }, required: ["city"] },
},
}],
});
console.log(resp.choices[0].message.tool_calls);Messages API
import os
import anthropic
client = anthropic.Anthropic(base_url="https://usecognia.xyz", api_key=os.environ["COGNIA_API_KEY"])
msg = client.messages.create(
model="ag/claude-opus-4-6-thinking",
max_tokens=512,
messages=[{"role": "user", "content": "Explain double-entry bookkeeping in two sentences."}],
)
print(msg.content[0].text)Errors
| Status | Code | Meaning |
|---|---|---|
| 400 | invalid_request, max_tokens_too_large, context_length_exceeded | Fix the request body. |
| 401 | invalid_api_key, revoked_api_key, expired_api_key | Check the key. |
| 402 | insufficient_balance | Add funds or lower max_tokens. |
| 403 | insufficient_scope, model_not_allowed, spend_limit_reached, request_limit_reached, account_frozen | Key or account restriction. |
| 404 | model_not_found | The model is not in the live catalog. |
| 409 / 422 | idempotency_in_progress, idempotency_key_reused | Idempotency conflict. |
| 429 | rate_limit_exceeded, concurrency_limit, upstream_rate_limited | Slow down. Honor Retry-After. |
| 502 / 503 / 504 | upstream_error, upstream_unreachable, upstream_no_capacity, model_temporarily_unavailable, upstream_timeout | The model request failed before returning output. You were not charged. |
Agent Keys and limits
- Scopes:
inferencefor completions and messages,models:readfor the model list. - Rate limit: requests per minute per key (default 60).
- Concurrency: simultaneous in-flight requests per key (default 4).
- Spend limit: optional daily, monthly, or lifetime cap in USD per key. A request is refused if the amount it would hold crosses the cap.
- Request limit: optional cap on requests per the same period. A request counts when it is accepted and sent to a model. Requests that are refused, or that fail and are released without charge, do not count.
- Revocation: a revoked key is rejected immediately. A request that was already running finishes and is billed normally.
- Model allowlist: optionally restrict a key to specific models.
- Idempotency: send
Idempotency-Keyon non-streaming requests to safely retry. A repeated key with the same body replays the stored response without a second charge for 24 hours.
Billing and usage
Each response includes x-cognia-request-id. Non-streaming responses also include x-cognia-cost-usd, and x-cognia-billing-state: pending_reconciliation when final usage is missing. A zero cost header while pending is not a final free charge. For streams, set stream_options.include_usage to receive the final usage chunk; the settled cost is in Usage. A stream you close early is settled only with verified usage; otherwise its funds remain held for reconciliation. Every balance movement is in the ledger. The earlier x-stovra-* header names and the stovra field in the model list are still sent for existing integrations.
Deposits
- Network: Robinhood Chain, chain ID 4663.
- Send only from a wallet you linked, and send the exact amount shown for that deposit. The amount identifies your deposit.
- Credits appear after 120 confirmations. Transfers that do not match a deposit are held for manual review, not lost.
- Stock Token deposits require an eligible residence and location, and a fresh quote. See the disclosures.
Chat workspace
The chat workspace uses the same catalog, prices, and ledger as the API. Chats are stored in your account until you delete them. Each message sends the earlier messages of that chat again as context, which counts as input tokens. Text and source files can be attached (up to 5 files, 200 KB each, 400 KB per message); they are sent as text and are never executed.
Scheduled agents
Create scheduled agents in the app. Runs are queued and executed by an isolated worker with the tools you grant. Each run records its steps, output, and cost. Scheduled agents run at most every 15 minutes and pause if your balance cannot cover a run.