AI Inference
Run AI inference through the Strongly AI Gateway. All inference endpoints proxy to the configured AI Gateway backend and support both synchronous and streaming responses. These endpoints are OpenAI-compatible, allowing drop-in replacement for existing OpenAI SDK integrations.
Responses from these endpoints are the AI Gateway's own response body and status, passed through as is: they are not wrapped in the standard data / meta envelope. The X-Request-Id header is still set.
POST /api/v1/ai-gateway/chat/completions
Create a chat completion
Generates a model response for the given conversation. Compatible with the OpenAI Chat Completions API format.
Scope: ai-gateway:inference
Headers Forwarded:
| Header | Description |
|---|---|
X-User-Id | Authenticated user ID |
X-Request-Id | Request trace ID |
X-Organization-ID | Organization context for multi-tenancy |
Authorization | Bearer token passed to the upstream provider |
Request Body:
{
"model": "gpt-4",
"messages": [
{ "role": "system", "content": "You are a helpful assistant." },
{ "role": "user", "content": "Explain Kubernetes in one sentence." }
],
"stream": false,
"max_tokens": 1024,
"temperature": 0.7,
"top_p": 1.0,
"stop": ["\n"]
}
| Field | Type | Required | Description |
|---|---|---|---|
model | string | Yes | Model ID or name registered in the AI Gateway |
messages | array | Yes | Conversation messages array |
messages[].role | string | Yes | Message role: system, user, or assistant |
messages[].content | string or array | Yes | Message content: text, or (on a user message) an array of content parts with images and PDF documents; see Images and PDF documents |
stream | boolean | No | Enable SSE streaming (default: false) |
max_tokens | integer | No | Maximum tokens to generate |
temperature | number | No | Sampling temperature (0.0 - 2.0) |
top_p | number | No | Nucleus sampling threshold (0.0 - 1.0) |
stop | string or string[] | No | Stop sequence(s) |
web_search_options | object | No | Search the web before answering. {} for the defaults; max_uses caps the searches. OpenAI, Anthropic and Gemini models use their vendor's own search; other models use Strongly's search when they can call tools. See Web Search. |
Images and PDF documents
A user message's content can be an array of parts, in the OpenAI Chat Completions shapes. The gateway gives each part to the model's vendor in that vendor's own input (OpenAI, Anthropic, Gemini and the other providers), streaming or not:
{
"role": "user",
"content": [
{ "type": "text", "text": "What does the chart on page 2 show?" },
{ "type": "image_url", "image_url": { "url": "data:image/png;base64,iVBORw0..." } },
{ "type": "file", "file": { "filename": "report.pdf", "file_data": "data:application/pdf;base64,JVBERi0..." } }
]
}
| Part | Shape | Taken by |
|---|---|---|
| Text | { "type": "text", "text": "..." } | Every chat model |
| Image | { "type": "image_url", "image_url": { "url": "data:<type>;base64,..." or "https://..." } }; PNG, JPEG, GIF and WebP | Models that declare image input (capabilityConfig.supportsVision) |
| PDF document | { "type": "file", "file": { "filename": "...", "file_data": "data:application/pdf;base64,..." } } | Models that declare PDF input (capabilityConfig.supportsPdf): the model reads the pages themselves (layout, tables, charts, scanned pages) |
A model's capabilityConfig comes from its certification and is shown with the model (GET /api/v1/ai-gateway/models/:id). A PDF sent to a model that does not declare PDF input is refused with 400 naming the model, and so is a part the model's vendor has no input for; a part is never dropped on the way. Send a PDF's text instead to a model that does not read PDFs. The document's pages count toward the request's input tokens in usage.
Response: 200 OK
{
"id": "chatcmpl-abc123",
"model": "gpt-4",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "Kubernetes is an open-source container orchestration platform."
},
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 25,
"completion_tokens": 12,
"total_tokens": 37
},
"created": 1706000000
}
POST /api/v1/ai-gateway/completions
Create a text completion
Generates a completion for the given prompt. Useful for non-conversational text generation tasks.
Scope: ai-gateway:inference
Headers Forwarded:
| Header | Description |
|---|---|
X-User-Id | Authenticated user ID |
X-Request-Id | Request trace ID |
X-Organization-ID | Organization context for multi-tenancy |
Authorization | Bearer token passed to the upstream provider |
Request Body:
{
"model": "gpt-6-luna",
"prompt": "Write a SQL query that selects all users where",
"stream": false,
"max_tokens": 256,
"temperature": 0.5
}
| Field | Type | Required | Description |
|---|---|---|---|
model | string | Yes | Model ID or name registered in the AI Gateway |
prompt | string | Yes | Text prompt to complete |
stream | boolean | No | Enable SSE streaming (default: false) |
max_tokens | integer | No | Maximum tokens to generate |
temperature | number | No | Sampling temperature (0.0 - 2.0) |
Response: 200 OK
{
"id": "cmpl-abc123",
"model": "gpt-6-luna",
"choices": [
{
"text": " status = 'active' ORDER BY created_at DESC;",
"index": 0,
"finish_reason": "stop",
"logprobs": null
}
],
"usage": {
"prompt_tokens": 12,
"completion_tokens": 14,
"total_tokens": 26
},
"created": 1706000000
}
POST /api/v1/ai-gateway/embeddings
Generate embeddings
Creates an embedding vector for the given input text. Supports single strings or batches of strings.
Scope: ai-gateway:inference
Request Body:
{
"model": "text-embedding-ada-002",
"input": "Kubernetes pod scheduling explained"
}
Or batch input:
{
"model": "text-embedding-ada-002",
"input": [
"Kubernetes pod scheduling explained",
"Docker container networking basics"
]
}
| Field | Type | Required | Description |
|---|---|---|---|
model | string | Yes | Embedding model ID or name registered in the AI Gateway |
input | string or string[] | Yes | Text to embed (single string or array of strings) |
Response: 200 OK
{
"data": [
{
"embedding": [0.0023064255, -0.009327292, 0.015797347, "..."],
"index": 0
}
],
"model": "text-embedding-ada-002",
"usage": {
"prompt_tokens": 6,
"total_tokens": 6
}
}
POST /api/v1/ai-gateway/tokenize
Count a prompt's tokens with a model's own tokenizer, and give that model's context length.
Self-hosted (vLLM-served) models only; any other model is refused with 400.
Scope: ai-gateway:inference
Request Body
| Field | Type | Required | Description |
|---|---|---|---|
model | string | Yes | The model |
prompt | string | Yes | The text to count |
Response 200 OK: the model server's answer, as it gives it: count, the tokens and
max_model_len when the server returns them.
GET /api/v1/ai-gateway/audio/speech/voices
The text-to-speech voices you can give POST /api/v1/ai-gateway/audio/speech, each with its
id, name, provider and gender.
Scope: ai-gateway:inference
| Parameter | Type | Required | Description |
|---|---|---|---|
provider | string | No | Only this provider's voices |
Response 200 OK
POST /api/v1/ai-gateway/generations
Start generating an image, a video or music with a generation model. The generation runs as a job: the response gives its id to follow.
Scope: ai-gateway:inference
Request Body
| Field | Type | Required | Description |
|---|---|---|---|
model | string | Yes | A generation model |
prompt | string | Yes | What to generate |
type | string | No | image, video or music (default from the model) |
Response 200 OK: the generation job, with its id and status.
DELETE /api/v1/ai-gateway/generations/:id
Cancel a generation job.
Scope: ai-gateway:inference
| Parameter | Type | Required | Description |
|---|---|---|---|
id | string | Yes | The generation job |
Response 200 OK
POST /api/v1/ai-gateway/generations/images
Generate images from a text prompt with an image generation model.
Scope: ai-gateway:inference
Request Body
| Field | Type | Required | Description |
|---|---|---|---|
model | string | Yes | An image generation model |
prompt | string | Yes | What to draw |
n | number | No | How many images |
size | string | No | The image size, e.g. 1024x1024 |
Response 200 OK: the model server's answer (OpenAI's image shape: data, each image a url or b64_json).
POST /api/v1/ai-gateway/generations/videos
Start generating a video from a text prompt. Video generation runs as a job: the response
gives its id; follow it with GET /api/v1/ai-gateway/generations/:id.
Scope: ai-gateway:inference
Request Body
| Field | Type | Required | Description |
|---|---|---|---|
model | string | Yes | A video generation model |
prompt | string | Yes | What to generate |
Response 200 OK: the generation job, with its id and status.
POST /api/v1/ai-gateway/generations/music
Start generating music from a text prompt. Music generation runs as a job, followed like a video's.
Scope: ai-gateway:inference
Request Body
| Field | Type | Required | Description |
|---|---|---|---|
model | string | Yes | A music generation model |
prompt | string | Yes | What to generate |
Response 200 OK: the generation job, with its id and status.
GET /api/v1/ai-gateway/generations/:id
A generation job's status and, once it is complete, its result.
Scope: ai-gateway:inference
| Parameter | Type | Required | Description |
|---|---|---|---|
id | string | Yes | The generation job |
Response 200 OK: the job: status and, when complete, where its output is.
POST /api/v1/ai-gateway/audio/transcriptions
Transcribe an audio file to text with a speech-to-text model. multipart/form-data, in
OpenAI's shape.
Scope: ai-gateway:inference
Form fields
| Field | Type | Required | Description |
|---|---|---|---|
model | string | Yes | A speech-to-text model |
file | file | Yes | The audio |
language | string | No | The spoken language's code |
prompt | string | No | Text that guides the transcription |
response_format | string | No | json (the default), text, srt, verbose_json or vtt |
temperature | number | No | Sampling temperature |
Response 200 OK: the transcription (text, or the format asked for).
POST /api/v1/ai-gateway/audio/translations
Translate an audio file's speech to English text. multipart/form-data with model and file,
as for transcriptions.
Scope: ai-gateway:inference
Response 200 OK: the English text.
POST /api/v1/ai-gateway/moderations
Check text with a moderation model.
Scope: ai-gateway:inference
Request Body
| Field | Type | Required | Description |
|---|---|---|---|
input | string | Yes | The text to check |
model | string | No | A moderation model |
Response 200 OK: OpenAI's moderation shape: results, each with flagged, categories and category_scores.
POST /api/v1/ai-gateway/rerank
Order documents by how relevant each is to a query, with a reranking model.
Scope: ai-gateway:inference
Request Body
| Field | Type | Required | Description |
|---|---|---|---|
model | string | Yes | A reranking model |
query | string | Yes | What to rank against |
documents | array | Yes | The documents (strings) |
Response 200 OK: the model's ranking in Cohere's shape: results, each a document's index and relevance_score, most relevant first.
SSE Streaming Format
When stream is set to true on the chat completions or text completions endpoints, the response uses Server-Sent Events (SSE) instead of returning a single JSON object.
Response Headers:
Content-Type: text/event-stream
Cache-Control: no-cache
Connection: keep-alive
Stream Chunks:
Each chunk is delivered as a data: event followed by a JSON object and two newlines:
data: {"id":"chatcmpl-abc123","model":"gpt-4","choices":[{"index":0,"delta":{"role":"assistant","content":"Kubernetes"},"finish_reason":null}],"created":1706000000}
data: {"id":"chatcmpl-abc123","model":"gpt-4","choices":[{"index":0,"delta":{"content":" is"},"finish_reason":null}],"created":1706000000}
data: {"id":"chatcmpl-abc123","model":"gpt-4","choices":[{"index":0,"delta":{"content":" an"},"finish_reason":null}],"created":1706000000}
data: {"id":"chatcmpl-abc123","model":"gpt-4","choices":[{"index":0,"delta":{},"finish_reason":"stop"}],"created":1706000000}
data: [DONE]
| Field | Description |
|---|---|
delta.role | Present only in the first chunk |
delta.content | Token content (may be empty in the final chunk) |
finish_reason | null during streaming, stop or length on the final content chunk |
[DONE] | Signals the end of the stream |
The chunks above are from chat completions. Text completion chunks carry the generated text in each choice's text field instead of delta.
When consuming the stream, concatenate all delta.content values to reconstruct the full response. The usage field is not included in streamed chunks.