Skip to content

Documentation

Cognia API

Agent Keys call the same models as the chat workspace through an OpenAI-compatible API at https://usecognia.xyz/v1. OpenAI SDKs work by changing the base URL and key. The supported request fields are listed below; anything else is rejected with a clear error instead of being ignored.

Quickstart

  1. Sign in with email or a wallet.
  2. Add funds on the Billing page with USDG on Robinhood Chain.
  3. Create a key on the Agent Keys page. It is shown once, so store it in a secret manager.
  4. Pick a model id from Models or GET /v1/models.
curl
export COGNIA_API_KEY="sk-cog-..."

curl https://usecognia.xyz/v1/chat/completions \
  -H "Authorization: Bearer $COGNIA_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "ag/claude-opus-4-6-thinking",
    "max_tokens": 300,
    "messages": [{"role": "user", "content": "Write a haiku about ledgers."}]
  }'
Python (openai>=1.0)
import os
from openai import OpenAI

client = OpenAI(base_url="https://usecognia.xyz/v1", api_key=os.environ["COGNIA_API_KEY"])

resp = client.chat.completions.create(
    model="ag/claude-opus-4-6-thinking",
    max_tokens=300,
    messages=[{"role": "user", "content": "Write a haiku about ledgers."}],
)
print(resp.choices[0].message.content)
Node.js (openai v4+)
import OpenAI from "openai";

const client = new OpenAI({ baseURL: "https://usecognia.xyz/v1", apiKey: process.env.COGNIA_API_KEY });

const resp = await client.chat.completions.create({
  model: "ag/claude-opus-4-6-thinking",
  max_tokens: 300,
  messages: [{ role: "user", content: "Write a haiku about ledgers." }],
});
console.log(resp.choices[0].message.content);

Authentication

Send your key as Authorization: Bearer sk-cog-.... The Messages endpoint also accepts x-api-key. Keys are stored only as keyed hashes, so a lost key cannot be recovered: revoke it and create a new one. Browser sessions are not accepted on /v1, and Agent Keys are not accepted by the web app's own endpoints. Never put a key in client-side code.

API reference

EndpointPurpose
GET /v1/modelsOpenAI-format model list. Each model has a cognia object with prices in USD per 1M tokens and capabilities. Key optional.
POST /v1/chat/completionsOpenAI Chat Completions. Supports stream, tools, tool_choice, response_format, stop, temperature, top_p, seed, max_tokens or max_completion_tokens. n must be 1.
POST /v1/messagesAnthropic Messages format for models marked "Messages API".
POST /v1/images/generationsNot available yet. Returns 501.

If you omit max_tokens, the model's default output limit applies, lowered if needed to what your balance and the key's budget can cover; the limit actually used is returned in x-cognia-max-tokens-applied. A max_tokens you set yourself is never changed: if it cannot be covered, the request is refused with 402. Output is always bounded so the amount held is a true maximum.

Streaming

Set "stream": true to receive Server-Sent Events. Pass stream_options: {"include_usage": true} to get the final usage chunk, exactly as with OpenAI.

Python streaming
stream = client.chat.completions.create(
    model="ag/claude-opus-4-6-thinking",
    stream=True,
    stream_options={"include_usage": True},
    messages=[{"role": "user", "content": "Count to five."}],
)
for chunk in stream:
    if chunk.choices and chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="")

Tool calls

Tool definitions and tool call deltas pass through unchanged for models that support tools (see the Tools badge). Your code executes the tool and sends the result back as a tool message.

Node.js tool call
const resp = await client.chat.completions.create({
  model: "ag/claude-opus-4-6-thinking",
  messages: [{ role: "user", content: "What's the weather in Jakarta?" }],
  tools: [{
    type: "function",
    function: {
      name: "get_weather",
      parameters: { type: "object", properties: { city: { type: "string" } }, required: ["city"] },
    },
  }],
});
console.log(resp.choices[0].message.tool_calls);

Messages API

Python (anthropic SDK)
import os
import anthropic

client = anthropic.Anthropic(base_url="https://usecognia.xyz", api_key=os.environ["COGNIA_API_KEY"])
msg = client.messages.create(
    model="ag/claude-opus-4-6-thinking",
    max_tokens=512,
    messages=[{"role": "user", "content": "Explain double-entry bookkeeping in two sentences."}],
)
print(msg.content[0].text)

Errors

StatusCodeMeaning
400invalid_request, max_tokens_too_large, context_length_exceededFix the request body.
401invalid_api_key, revoked_api_key, expired_api_keyCheck the key.
402insufficient_balanceAdd funds or lower max_tokens.
403insufficient_scope, model_not_allowed, spend_limit_reached, request_limit_reached, account_frozenKey or account restriction.
404model_not_foundThe model is not in the live catalog.
409 / 422idempotency_in_progress, idempotency_key_reusedIdempotency conflict.
429rate_limit_exceeded, concurrency_limit, upstream_rate_limitedSlow down. Honor Retry-After.
502 / 503 / 504upstream_error, upstream_unreachable, upstream_no_capacity, model_temporarily_unavailable, upstream_timeoutThe model request failed before returning output. You were not charged.

Agent Keys and limits

  • Scopes: inference for completions and messages, models:read for the model list.
  • Rate limit: requests per minute per key (default 60).
  • Concurrency: simultaneous in-flight requests per key (default 4).
  • Spend limit: optional daily, monthly, or lifetime cap in USD per key. A request is refused if the amount it would hold crosses the cap.
  • Request limit: optional cap on requests per the same period. A request counts when it is accepted and sent to a model. Requests that are refused, or that fail and are released without charge, do not count.
  • Revocation: a revoked key is rejected immediately. A request that was already running finishes and is billed normally.
  • Model allowlist: optionally restrict a key to specific models.
  • Idempotency: send Idempotency-Key on non-streaming requests to safely retry. A repeated key with the same body replays the stored response without a second charge for 24 hours.

Billing and usage

Each response includes x-cognia-request-id. Non-streaming responses also include x-cognia-cost-usd, and x-cognia-billing-state: pending_reconciliation when final usage is missing. A zero cost header while pending is not a final free charge. For streams, set stream_options.include_usage to receive the final usage chunk; the settled cost is in Usage. A stream you close early is settled only with verified usage; otherwise its funds remain held for reconciliation. Every balance movement is in the ledger. The earlier x-stovra-* header names and the stovra field in the model list are still sent for existing integrations.

Deposits

  • Network: Robinhood Chain, chain ID 4663.
  • Send only from a wallet you linked, and send the exact amount shown for that deposit. The amount identifies your deposit.
  • Credits appear after 120 confirmations. Transfers that do not match a deposit are held for manual review, not lost.
  • Stock Token deposits require an eligible residence and location, and a fresh quote. See the disclosures.

Chat workspace

The chat workspace uses the same catalog, prices, and ledger as the API. Chats are stored in your account until you delete them. Each message sends the earlier messages of that chat again as context, which counts as input tokens. Text and source files can be attached (up to 5 files, 200 KB each, 400 KB per message); they are sent as text and are never executed.

Scheduled agents

Create scheduled agents in the app. Runs are queued and executed by an isolated worker with the tools you grant. Each run records its steps, output, and cost. Scheduled agents run at most every 15 minutes and pause if your balance cannot cover a run.