Chat Completions

OpenAI-compatible POST /v1/chat/completions including SSE streaming.

Endpoint

POST /v1/chat/completions

OpenAI-compatible Chat Completions, including SSE streaming.

Request body

FieldRequiredNotes
modelYesFabric Model slug allowed for your Organisation
messagesYesOpenAI-style role/content array
streamNotrue for SSE chunks
max_tokensNoAffects Reserve estimation; defaults from Fabric Model
{
  "model": "apparatus-pro",
  "messages": [
    {"role": "system", "content": "Be concise."},
    {"role": "user", "content": "Summarise Fabrix."}
  ],
  "stream": false
}

Response

Successful responses follow the OpenAI chat completion shape. The model field always echoes the Fabric Model slug. Upstream provider headers and vendor error text are stripped.

Streaming

Set stream: true. The gateway emits Server-Sent Events in OpenAI chunk form. During streaming, Fabrix best-effort stops generation when running cost approaches the Reserve.

Reserve and billing

Before any provider call, Fabrix estimates tokens and holds a Reserve against the Organisation Balance (and Project Budget when set). Unused Reserve is released; actual usage is then debited. If a Reserve cannot be held, the request fails with 402 before upstream traffic.

Organisations in internal billing mode skip wallet debit but still record usage and provider cost for analytics.