Skip to main content

AI Inference

Run AI inference through the Strongly AI Gateway. All inference endpoints proxy to the configured AI Gateway backend and support both synchronous and streaming responses. These endpoints are OpenAI-compatible, allowing drop-in replacement for existing OpenAI SDK integrations.

Responses from these endpoints are the AI Gateway's own response body and status, passed through as is: they are not wrapped in the standard data / meta envelope. The X-Request-Id header is still set.


POST /api/v1/ai-gateway/chat/completions​

Create a chat completion

Generates a model response for the given conversation. Compatible with the OpenAI Chat Completions API format.

Scope: ai-gateway:inference

Headers Forwarded:

HeaderDescription
X-User-IdAuthenticated user ID
X-Request-IdRequest trace ID
X-Organization-IDOrganization context for multi-tenancy
AuthorizationBearer token passed to the upstream provider

Request Body:

{
"model": "gpt-4",
"messages": [
{ "role": "system", "content": "You are a helpful assistant." },
{ "role": "user", "content": "Explain Kubernetes in one sentence." }
],
"stream": false,
"max_tokens": 1024,
"temperature": 0.7,
"top_p": 1.0,
"stop": ["\n"]
}
FieldTypeRequiredDescription
modelstringYesModel ID or name registered in the AI Gateway
messagesarrayYesConversation messages array
messages[].rolestringYesMessage role: system, user, or assistant
messages[].contentstring or arrayYesMessage content: text, or (on a user message) an array of content parts with images and PDF documents; see Images and PDF documents
streambooleanNoEnable SSE streaming (default: false)
max_tokensintegerNoMaximum tokens to generate
temperaturenumberNoSampling temperature (0.0 - 2.0)
top_pnumberNoNucleus sampling threshold (0.0 - 1.0)
stopstring or string[]NoStop sequence(s)
web_search_optionsobjectNoSearch the web before answering. {} for the defaults; max_uses caps the searches. OpenAI, Anthropic and Gemini models use their vendor's own search; other models use Strongly's search when they can call tools. See Web Search.

Images and PDF documents​

A user message's content can be an array of parts, in the OpenAI Chat Completions shapes. The gateway gives each part to the model's vendor in that vendor's own input (OpenAI, Anthropic, Gemini and the other providers), streaming or not:

{
"role": "user",
"content": [
{ "type": "text", "text": "What does the chart on page 2 show?" },
{ "type": "image_url", "image_url": { "url": "data:image/png;base64,iVBORw0..." } },
{ "type": "file", "file": { "filename": "report.pdf", "file_data": "data:application/pdf;base64,JVBERi0..." } }
]
}
PartShapeTaken by
Text{ "type": "text", "text": "..." }Every chat model
Image{ "type": "image_url", "image_url": { "url": "data:<type>;base64,..." or "https://..." } }; PNG, JPEG, GIF and WebPModels that declare image input (capabilityConfig.supportsVision)
PDF document{ "type": "file", "file": { "filename": "...", "file_data": "data:application/pdf;base64,..." } }Models that declare PDF input (capabilityConfig.supportsPdf): the model reads the pages themselves (layout, tables, charts, scanned pages)

A model's capabilityConfig comes from its certification and is shown with the model (GET /api/v1/ai-gateway/models/:id). A PDF sent to a model that does not declare PDF input is refused with 400 naming the model, and so is a part the model's vendor has no input for; a part is never dropped on the way. Send a PDF's text instead to a model that does not read PDFs. The document's pages count toward the request's input tokens in usage.

Response: 200 OK

{
"id": "chatcmpl-abc123",
"model": "gpt-4",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "Kubernetes is an open-source container orchestration platform."
},
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 25,
"completion_tokens": 12,
"total_tokens": 37
},
"created": 1706000000
}

POST /api/v1/ai-gateway/completions​

Create a text completion

Generates a completion for the given prompt. Useful for non-conversational text generation tasks.

Scope: ai-gateway:inference

Headers Forwarded:

HeaderDescription
X-User-IdAuthenticated user ID
X-Request-IdRequest trace ID
X-Organization-IDOrganization context for multi-tenancy
AuthorizationBearer token passed to the upstream provider

Request Body:

{
"model": "gpt-6-luna",
"prompt": "Write a SQL query that selects all users where",
"stream": false,
"max_tokens": 256,
"temperature": 0.5
}
FieldTypeRequiredDescription
modelstringYesModel ID or name registered in the AI Gateway
promptstringYesText prompt to complete
streambooleanNoEnable SSE streaming (default: false)
max_tokensintegerNoMaximum tokens to generate
temperaturenumberNoSampling temperature (0.0 - 2.0)

Response: 200 OK

{
"id": "cmpl-abc123",
"model": "gpt-6-luna",
"choices": [
{
"text": " status = 'active' ORDER BY created_at DESC;",
"index": 0,
"finish_reason": "stop",
"logprobs": null
}
],
"usage": {
"prompt_tokens": 12,
"completion_tokens": 14,
"total_tokens": 26
},
"created": 1706000000
}

POST /api/v1/ai-gateway/embeddings​

Generate embeddings

Creates an embedding vector for the given input text. Supports single strings or batches of strings.

Scope: ai-gateway:inference

Request Body:

{
"model": "text-embedding-ada-002",
"input": "Kubernetes pod scheduling explained"
}

Or batch input:

{
"model": "text-embedding-ada-002",
"input": [
"Kubernetes pod scheduling explained",
"Docker container networking basics"
]
}
FieldTypeRequiredDescription
modelstringYesEmbedding model ID or name registered in the AI Gateway
inputstring or string[]YesText to embed (single string or array of strings)

Response: 200 OK

{
"data": [
{
"embedding": [0.0023064255, -0.009327292, 0.015797347, "..."],
"index": 0
}
],
"model": "text-embedding-ada-002",
"usage": {
"prompt_tokens": 6,
"total_tokens": 6
}
}

POST /api/v1/ai-gateway/tokenize​

Count a prompt's tokens with a model's own tokenizer, and give that model's context length. Self-hosted (vLLM-served) models only; any other model is refused with 400.

Scope: ai-gateway:inference

Request Body

FieldTypeRequiredDescription
modelstringYesThe model
promptstringYesThe text to count

Response 200 OK: the model server's answer, as it gives it: count, the tokens and max_model_len when the server returns them.


GET /api/v1/ai-gateway/audio/speech/voices​

The text-to-speech voices you can give POST /api/v1/ai-gateway/audio/speech, each with its id, name, provider and gender.

Scope: ai-gateway:inference

ParameterTypeRequiredDescription
providerstringNoOnly this provider's voices

Response 200 OK


POST /api/v1/ai-gateway/generations​

Start generating an image, a video or music with a generation model. The generation runs as a job: the response gives its id to follow.

Scope: ai-gateway:inference

Request Body

FieldTypeRequiredDescription
modelstringYesA generation model
promptstringYesWhat to generate
typestringNoimage, video or music (default from the model)

Response 200 OK: the generation job, with its id and status.


DELETE /api/v1/ai-gateway/generations/:id​

Cancel a generation job.

Scope: ai-gateway:inference

ParameterTypeRequiredDescription
idstringYesThe generation job

Response 200 OK


POST /api/v1/ai-gateway/generations/images​

Generate images from a text prompt with an image generation model.

Scope: ai-gateway:inference

Request Body

FieldTypeRequiredDescription
modelstringYesAn image generation model
promptstringYesWhat to draw
nnumberNoHow many images
sizestringNoThe image size, e.g. 1024x1024

Response 200 OK: the model server's answer (OpenAI's image shape: data, each image a url or b64_json).


POST /api/v1/ai-gateway/generations/videos​

Start generating a video from a text prompt. Video generation runs as a job: the response gives its id; follow it with GET /api/v1/ai-gateway/generations/:id.

Scope: ai-gateway:inference

Request Body

FieldTypeRequiredDescription
modelstringYesA video generation model
promptstringYesWhat to generate

Response 200 OK: the generation job, with its id and status.


POST /api/v1/ai-gateway/generations/music​

Start generating music from a text prompt. Music generation runs as a job, followed like a video's.

Scope: ai-gateway:inference

Request Body

FieldTypeRequiredDescription
modelstringYesA music generation model
promptstringYesWhat to generate

Response 200 OK: the generation job, with its id and status.


GET /api/v1/ai-gateway/generations/:id​

A generation job's status and, once it is complete, its result.

Scope: ai-gateway:inference

ParameterTypeRequiredDescription
idstringYesThe generation job

Response 200 OK: the job: status and, when complete, where its output is.


POST /api/v1/ai-gateway/audio/transcriptions​

Transcribe an audio file to text with a speech-to-text model. multipart/form-data, in OpenAI's shape.

Scope: ai-gateway:inference

Form fields

FieldTypeRequiredDescription
modelstringYesA speech-to-text model
filefileYesThe audio
languagestringNoThe spoken language's code
promptstringNoText that guides the transcription
response_formatstringNojson (the default), text, srt, verbose_json or vtt
temperaturenumberNoSampling temperature

Response 200 OK: the transcription (text, or the format asked for).


POST /api/v1/ai-gateway/audio/translations​

Translate an audio file's speech to English text. multipart/form-data with model and file, as for transcriptions.

Scope: ai-gateway:inference

Response 200 OK: the English text.


POST /api/v1/ai-gateway/moderations​

Check text with a moderation model.

Scope: ai-gateway:inference

Request Body

FieldTypeRequiredDescription
inputstringYesThe text to check
modelstringNoA moderation model

Response 200 OK: OpenAI's moderation shape: results, each with flagged, categories and category_scores.


POST /api/v1/ai-gateway/rerank​

Order documents by how relevant each is to a query, with a reranking model.

Scope: ai-gateway:inference

Request Body

FieldTypeRequiredDescription
modelstringYesA reranking model
querystringYesWhat to rank against
documentsarrayYesThe documents (strings)

Response 200 OK: the model's ranking in Cohere's shape: results, each a document's index and relevance_score, most relevant first.


SSE Streaming Format​

When stream is set to true on the chat completions or text completions endpoints, the response uses Server-Sent Events (SSE) instead of returning a single JSON object.

Response Headers:

Content-Type: text/event-stream
Cache-Control: no-cache
Connection: keep-alive

Stream Chunks:

Each chunk is delivered as a data: event followed by a JSON object and two newlines:

data: {"id":"chatcmpl-abc123","model":"gpt-4","choices":[{"index":0,"delta":{"role":"assistant","content":"Kubernetes"},"finish_reason":null}],"created":1706000000}

data: {"id":"chatcmpl-abc123","model":"gpt-4","choices":[{"index":0,"delta":{"content":" is"},"finish_reason":null}],"created":1706000000}

data: {"id":"chatcmpl-abc123","model":"gpt-4","choices":[{"index":0,"delta":{"content":" an"},"finish_reason":null}],"created":1706000000}

data: {"id":"chatcmpl-abc123","model":"gpt-4","choices":[{"index":0,"delta":{},"finish_reason":"stop"}],"created":1706000000}

data: [DONE]

FieldDescription
delta.rolePresent only in the first chunk
delta.contentToken content (may be empty in the final chunk)
finish_reasonnull during streaming, stop or length on the final content chunk
[DONE]Signals the end of the stream

The chunks above are from chat completions. Text completion chunks carry the generated text in each choice's text field instead of delta.

tip

When consuming the stream, concatenate all delta.content values to reconstruct the full response. The usage field is not included in streamed chunks.