Skip to main content
This endpoint follows the OpenAI Chat Completions shape. Streaming returns data-framed completion chunks followed by data: [DONE].
The request must contain at least one message. The gateway validates message size, generation limits, sampling values, tools, and structured-output settings before sending the request.

Request fields

Response

A non-streaming call returns a choices array, where each choice has an index, a message, and a finish_reason, plus a usage object with prompt_tokens, completion_tokens, and total_tokens.

Function calling

Describe your functions in tools, let the model decide with tool_choice: "auto", run the function yourself, and send the result back in a tool message.

Behavior to know

  • Cancelling the HTTP request cancels the upstream generation.
  • This endpoint does not create asynchronous jobs, so callbacks and webhooks are not accepted.
  • Possible errors are 400, 401, 403, 404, 413, 429, 502, 503, and 504. See errors and retries.
For field-by-field behavior, see compatibility and streaming.