Skip to main content

API Endpoints

You call the AI Gateway in one of two ways:

  • From outside the platform (your laptop, a CI job, another service): through the Strongly REST API at https://<your-platform>/api/v1/ai-gateway/..., with an API key.
  • From inside the platform (an app, a workspace, a workflow node, a job): through the AI Gateway address in STRONGLY_SERVICES. The platform authenticates these calls for you; no key is needed.

Both speak the OpenAI API format, so OpenAI client libraries work against either.

From outside the platform: the REST API​

Authentication​

Send an API key in the X-API-Key header. Create one in Profile > Security > API Keys (see Authentication). Inference needs the ai-gateway:inference scope; managing models needs ai-gateway:read and ai-gateway:write.

curl -X POST https://<your-platform>/api/v1/ai-gateway/chat/completions \
-H "X-API-Key: $STRONGLY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "<model id>",
"messages": [{"role": "user", "content": "Explain Kubernetes in one sentence."}]
}'

model is the ID of a model you can use: list them with GET /api/v1/ai-gateway/models (see AI Models).

With an OpenAI client​

Point the client at the REST API and send your key in X-API-Key:

from openai import OpenAI

client = OpenAI(
base_url="https://<your-platform>/api/v1/ai-gateway",
api_key="unused", # the platform reads X-API-Key, not the OpenAI key
default_headers={"X-API-Key": "<your Strongly API key>"},
)

reply = client.chat.completions.create(
model="<model id>",
messages=[{"role": "user", "content": "Hello"}],
)
print(reply.choices[0].message.content)

The Python SDK wraps the same routes with the key already set.

Inference routes​

RouteWhat it does
POST /api/v1/ai-gateway/chat/completionsChat completion; "stream": true streams it as server-sent events
POST /api/v1/ai-gateway/completionsText completion
POST /api/v1/ai-gateway/embeddingsEmbeddings
POST /api/v1/ai-gateway/tokenizeCount a text's tokens for a model
POST /api/v1/ai-gateway/audio/speechText to speech
GET /api/v1/ai-gateway/audio/speech/voicesA speech model's voices
POST /api/v1/ai-gateway/audio/transcriptionsSpeech to text (multipart upload)
POST /api/v1/ai-gateway/audio/translationsSpeech to English text (multipart upload)
POST /api/v1/ai-gateway/generationsStart a generation of any type
POST /api/v1/ai-gateway/generations/imagesGenerate images
POST /api/v1/ai-gateway/generations/videosGenerate a video
POST /api/v1/ai-gateway/generations/musicGenerate music
GET /api/v1/ai-gateway/generations/:idA generation's status and result
DELETE /api/v1/ai-gateway/generations/:idCancel a generation
POST /api/v1/ai-gateway/moderationsModerate content
POST /api/v1/ai-gateway/rerankRerank documents against a query

Each route's request and response are in the AI Inference reference.

These routes answer with the AI Gateway's own response body and status, as an OpenAI client expects them: they are not wrapped in the REST API's { data, meta } envelope. The X-Request-Id header is still set, so you can trace a request.

Managing models, keys and usage​

Everything else about the gateway is a regular REST resource with the { data, meta } envelope:

AreaReference
Models: list, add, deploy, start, stop, resize, settings, metrics, logsAI Models
Your providers' API keysProvider Keys
Model routers (one name, several models)Routers
Guardrails on a modelGuardrails
Usage, latency and cost analyticsAI Analytics
Fine-tuning jobsFine-Tuning
Data Forge (training data)Data Forge

Errors​

Inference routes pass the gateway's error body and status through. A request the REST API itself refuses (no key, a key without the scope, a missing model) answers an RFC 9457 problem with type, title, status and detail:

StatusMeaning
401No API key, or not a valid one
403The key lacks the scope, or you may not use the model
404No such model (or one you cannot see)
429Rate limited: retry after the Retry-After header
502The AI Gateway could not be reached

From inside the platform: STRONGLY_SERVICES​

Apps, workspaces, jobs and workflow nodes get the platform's services in the STRONGLY_SERVICES environment variable. Its aigateway entry carries the gateway's in-cluster base_url and the models you chose for the workload (available_models). Call the gateway's OpenAI-compatible /v1/... routes at that address and name the model with its Strongly ID in X-Model-Id:

import json, os, requests

services = json.loads(os.environ["STRONGLY_SERVICES"])
gateway = services["services"]["aigateway"]
model = gateway["available_models"][0]

reply = requests.post(
f"{gateway['base_url']}/v1/chat/completions",
headers={"Content-Type": "application/json", "X-Model-Id": model["_id"]},
json={
"model": model["vendor_model_id"],
"messages": [{"role": "user", "content": "Hello"}],
},
)
print(reply.json()["choices"][0]["message"]["content"])

The platform authenticates these calls as the workload's owner, so they count toward that user's usage and budgets like their own requests. See Platform Services for every entry in STRONGLY_SERVICES and examples in Node.js.