Chat Completions
OpenAI-compatible POST /v1/chat/completions including SSE streaming.
Endpoint
POST /v1/chat/completionsOpenAI-compatible Chat Completions, including SSE streaming.
Request body
| Field | Required | Notes |
|---|---|---|
model | Yes | Fabric Model slug allowed for your Organisation |
messages | Yes | OpenAI-style role/content array |
stream | No | true for SSE chunks |
max_tokens | No | Affects Reserve estimation; defaults from Fabric Model |
{
"model": "apparatus-pro",
"messages": [
{"role": "system", "content": "Be concise."},
{"role": "user", "content": "Summarise Fabrix."}
],
"stream": false
}Response
Successful responses follow the OpenAI chat completion shape. The model field always echoes the Fabric Model slug. Upstream provider headers and vendor error text are stripped.
Streaming
Set stream: true. The gateway emits Server-Sent Events in OpenAI chunk form. During streaming, Fabrix best-effort stops generation when running cost approaches the Reserve.
Reserve and billing
Before any provider call, Fabrix estimates tokens and holds a Reserve against the Organisation Balance (and Project Budget when set). Unused Reserve is released; actual usage is then debited. If a Reserve cannot be held, the request fails with 402 before upstream traffic.
Organisations in internal billing mode skip wallet debit but still record usage and provider cost for analytics.