Skip to content
Testnet demo on Whitechain Sepolia: test tokens (ITC) only, no real money. Requests are served by a real open-weight model on the founder's machine when it is online, otherwise by a demo seller that returns canned responses.
Inferit

API reference

Base URL: https://api-production-c74b9.up.railway.app (this site's current API). Inference is OpenAI-compatible; everything else is plain JSON.

Authentication

  • Inference takes an API key: Authorization: Bearer ik_… or x-api-key: ik_…. Without either, it takes an x402 payment instead (below).
  • Account routes (keys, balance, usage, seller) take the session token from sign-in as Authorization: Bearer <session>. Sessions last 24h. Balance and usage also accept your API key.
  • Public routes (models, prices, marketplace, stats, rails) need no auth and are cacheable.

Money

All amounts are integers in millionths of a unit (fields ending in Micro), serialised as decimal strings. On the testnet the unit is ITC, which counts as one test dollar and has no monetary value: "1500000" = 1.50 ITC. Prices are in the same micro-units per 1M tokens, quoted in $ so they compare with list prices, and are all-in (seller price + platform fee) unless a field says otherwise. Percentages are numbers with two decimals; a discount is rounded down so it is never overstated.

Inference

POST/v1/chat/completionsAPI key
OpenAI Chat Completions. Supports stream: true (SSE relayed byte for byte). Before routing, your available balance on the selected rail must cover the worst case: input estimate plus the output limit at the top-ranked offer. The output limit is the smaller of max_tokens and max_completion_tokens, or 1024 when you send neither, and the seller is held to it: a long answer without a limit stops at 1024 tokens with finish_reason: "length". You are then charged the metered cost.
example (model with a live offer)
curl https://api-production-c74b9.up.railway.app/v1/chat/completions \
  -H "Authorization: Bearer $INFERIT_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"qwen/qwen-2.5-14b-instruct","messages":[{"role":"user","content":"Hello"}]}'
POST/min{N}/v1/chat/completionsAPI key
Same, but only routes to offers at least N% below list (0–100). Also available as the x-min-discount header or a per-key setting; the strictest wins.
POST/anthropic/v1/messages—
Not implemented yet: returns 501 with a docs link. See Claude Code.

Response headers on every inference call:

HeaderMeaning
x-request-idRequest id, also in your usage log
x-inferit-attemptsOffers tried (failover happens only before the first byte)
x-inferit-buyer-cost-microWhat this request cost you, in micro-units (µITC on the testnet). A stream's final cost is known only after its last chunk; see your usage log.
x-inferit-offerOpaque hash of the offer that served you, not the seller's identity. On Whitechain the settlement itself is public; see On-chain visibility
x-inferit-railRail the request is funded and settled on
x-inferit-networkSettlement network, on every response: whitechain-sepolia today (whitechain-mainnet after launch)
x-inferit-testnet1 on every response from a testnet deployment: balances are test tokens with no value

Errors

Errors use OpenAI's shape. Notable codes: 402 insufficient_funds (with a key; without one, a 402 is an x402 offer), 404 model_not_found, 503 no_available_offers (with retry-after), 429 rate_limit_exceeded (per client IP, with retry-after; sign-in and dev routes have a stricter budget), 401 for a missing or bad key.

json
{
  "error": {
    "message": "Your available balance on whitechain does not cover the worst-case cost of this request.",
    "type": "insufficient_funds",
    "code": "insufficient_funds"
  }
}

x402: pay per request, no key

Standard x402 v2, exact scheme, settled into the escrow and credited to the payer. The walkthrough and client code are on Pay per request with x402.

POST/v1/chat/completionsnone → 402
With no Authorization or x-api-key header: 402 with a PAYMENT-REQUIRED header (base64 JSON) and the same document as the body. It offers one requirement: scheme: "exact", network: "eip155:<chainId>", asset = the token, payTo = the escrow, amount = the worst-case all-in cost of this exact body (at least 0.01 ITC), maxTimeoutSeconds and extra: { name, version }, the token's EIP-712 domain. The quote is pinned to the body for 120 s. The same applies to /min{N}/v1/chat/completions.
POST/v1/chat/completionsPAYMENT-SIGNATURE
The same body with a PAYMENT-SIGNATURE header (base64 JSON payload with an EIP-3009 authorization; X-PAYMENT is accepted too). The API verifies it, settles it first with depositWithAuthorization, then serves the request against the payer's escrow balance (streaming included). The response carries PAYMENT-RESPONSE, base64 of { success: true, transaction, network, payer }, plus the usual x-inferit-* headers. A payment that does not verify gets a fresh 402 with the reason in error. If the request fails after the payment settled, the credit stays in the payer's escrow balance and the error says so (an error.x402 object with credited, payer, amount and transaction; a funding 402 at that point becomes 503, so the client does not pay again). 503 x402_settlement_pending: the deposit was sent but its receipt did not arrive in time; if it lands it is the payer's escrow balance. 503 x402_unavailable: nothing was sent and nothing was charged.
POST/v1/faucetnone
Testnets only. Body { address }: the API mints 1,000 test ITC to that address with faucetTo and pays the gas. At most once per 24 hours per address (429 faucet_cooldown); rate-limited per IP. 404 when the server does not sponsor the faucet.
GET/.well-known/x402
Discovery: the paid resources and the requirement a client signs for (scheme, network, asset, payTo, extra), the minimum payment, the facilitator address and the faucet. 404 when x402 is off.

Marketplace (public)

GET/v1/models
OpenRouter-shaped model rows plus pricing (list price) and best_price (best all-in offer).
GET/v1/prices
List and best all-in price per model.
GET/api/marketplace
Per model: best all-in input/output price, discount vs list, sellers, healthy sellers, 24h requests and volume, uptime, TTFT p50.
GET/api/markets/:model
Aggregated order book: price levels with offer count, healthy count and the combined daily caps of the offers at each level. Never seller identities, endpoints or per-offer volume. URL-encode the model id.
GET/api/stats
Rolling 24h totals: requests, spend, savings vs list, tokens, offers, sellers.
GET/v1/rails
Settlement on Whitechain: network, testnet flag, chain id, escrow and token addresses, operator, owner, guardian, escrow version, health.
GET/v1/providers/resale-allowlist
Providers whose terms permit API-key resale, each with an evidence URL. Empty by default.
GET/health
Liveness, the build's git SHA, the network, chain id and testnet flag. Answers 503 when the database or the settlement rail is unhealthy (RPC down, escrow paused, operator low on gas).
GET/metrics
JSON operations metrics: settlement backlog, batches awaiting finality, chain head lag, operator gas runway, circuit-breaker headroom. May require a bearer token on hosted deployments.
GET/llms.txt
Machine-readable summary for agents, including the x402 and API-key onboarding paths.
example (annotated)
GET /api/marketplace
{
  "feeBps": 500,
  "markets": [{
    "modelId": "qwen/qwen-2.5-14b-instruct",
    "name": "Qwen: Qwen2.5 14B Instruct",
    "openWeight": true,
    "licenseId": "apache-2.0",
    "listInputPerM": "100000",        // µUSD per 1M tokens ($0.10)
    "listOutputPerM": "200000",
    "bestInputPerM": "84000",         // all-in, healthy offers only
    "bestOutputPerM": "168000",
    "bestDiscountPct": 16,
    "sellers": 3, "healthySellers": 2,
    "requests24h": 1840, "volume24hMicro": "2214551",
    "uptime24hPct": 99.2, "ttftP50Ms": 310
  }]
}
example (annotated)
GET /api/markets/qwen%2Fqwen-2.5-14b-instruct
{
  "modelId": "qwen/qwen-2.5-14b-instruct",
  "summary": { ...same shape as a marketplace row... },
  "levels": [
    { "inputPerM": "84000", "outputPerM": "168000", "offers": 2, "healthyOffers": 2, "capacityMicro": "18000000" },
    { "inputPerM": "95000", "outputPerM": "190000", "offers": 1, "healthyOffers": 0, "capacityMicro": null }
  ]
}

Sign-in

POST/v1/auth/evm/challenge
Body { address }. Returns { message, nonce }: a one-time message to sign with the wallet.
POST/v1/auth/evm/verify
Body { message, signature }. Verifies the signature and returns { token, expiresAt, account }.

Keys

POST/v1/keyssession
Body { name?, spendLimitMicro?, minDiscountPct? }. Returns the key once in key.
GET/v1/keyssession
Your keys (prefix only, never the secret).
DELETE/v1/keys/:idsession
Revoke immediately.

Buyer

GET/v1/balancesession or API key
Available, cap, spent, deposit, expiry, unsettled charges (pendingMicro) and usage settled on-chain but not yet final (pendingFinalityMicro). Available subtracts both.
PUT/v1/buyer/railsession
Body { rail: "whitechain" }. Selects Whitechain settlement for future requests ("credit" exists only on local development APIs).
GET/v1/usage?limit&cursorsession or API key
Per-request token counts, costs, latency, rail and settlement status. Paginate with nextCursor. Each row carries finality: pending until its settlement batch is finalized on-chain, then finalized. Rows never carry the seller's identity, a transaction hash or a batch id; your own settlements are listed with explorer links on your dashboard.
GET/v1/usage/export.csvsession or API key
The same as CSV. Never contains prompts or completions.

Seller

POST/v1/seller/offerssession
Body { model, kind: "endpoint" | "upstream_key", endpointUrl, authToken?, upstreamProvider?, upstreamKey?, pricing, capDailyMicro?, licenseAck }. pricing is { mode: "per_token", inputPerM, outputPerM, cacheReadPerM? } or { mode: "multiplier", multiplierBps } (10,000 = list price). The API checks the licence, the resale allowlist and an SSRF guard, sends a live 1-token probe, and requires you to be registered on your rail.
GET/v1/seller/offerssession
Your offers, including your own endpoint URLs and health.
PATCH/v1/seller/offers/:idsession
Pause/resume, reprice, change cap or endpoint.
DELETE/v1/seller/offers/:idsession
Delist; stored secrets are destroyed (hard delete).
GET/v1/seller/earningssession
Settled, pending and withdrawable earnings, with the settled part split into pendingFinalityMicro and finalizedMicro.