API Endpoints
You call the AI Gateway in one of two ways:
- From outside the platform (your laptop, a CI job, another service): through the Strongly REST API at
https://<your-platform>/api/v1/ai-gateway/..., with an API key. - From inside the platform (an app, a workspace, a workflow node, a job): through the AI Gateway address in
STRONGLY_SERVICES. The platform authenticates these calls for you; no key is needed.
Both speak the OpenAI API format, so OpenAI client libraries work against either.
From outside the platform: the REST API
Authentication
Send an API key in the X-API-Key header. Create one in Profile > Security > API Keys (see Authentication). Inference needs the ai-gateway:inference scope; managing models needs ai-gateway:read and ai-gateway:write.
curl -X POST https://<your-platform>/api/v1/ai-gateway/chat/completions \
-H "X-API-Key: $STRONGLY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "<model id>",
"messages": [{"role": "user", "content": "Explain Kubernetes in one sentence."}]
}'
model is the ID of a model you can use: list them with GET /api/v1/ai-gateway/models (see AI Models).
With an OpenAI client
Point the client at the REST API and send your key in X-API-Key:
from openai import OpenAI
client = OpenAI(
base_url="https://<your-platform>/api/v1/ai-gateway",
api_key="unused", # the platform reads X-API-Key, not the OpenAI key
default_headers={"X-API-Key": "<your Strongly API key>"},
)
reply = client.chat.completions.create(
model="<model id>",
messages=[{"role": "user", "content": "Hello"}],
)
print(reply.choices[0].message.content)
The Python SDK wraps the same routes with the key already set.
Inference routes
| Route | What it does |
|---|---|
POST /api/v1/ai-gateway/chat/completions | Chat completion; "stream": true streams it as server-sent events |
POST /api/v1/ai-gateway/completions | Text completion |
POST /api/v1/ai-gateway/embeddings | Embeddings |
POST /api/v1/ai-gateway/tokenize | Count a text's tokens for a model |
POST /api/v1/ai-gateway/audio/speech | Text to speech |
GET /api/v1/ai-gateway/audio/speech/voices | A speech model's voices |
POST /api/v1/ai-gateway/audio/transcriptions | Speech to text (multipart upload) |
POST /api/v1/ai-gateway/audio/translations | Speech to English text (multipart upload) |
POST /api/v1/ai-gateway/generations | Start a generation of any type |
POST /api/v1/ai-gateway/generations/images | Generate images |
POST /api/v1/ai-gateway/generations/videos | Generate a video |
POST /api/v1/ai-gateway/generations/music | Generate music |
GET /api/v1/ai-gateway/generations/:id | A generation's status and result |
DELETE /api/v1/ai-gateway/generations/:id | Cancel a generation |
POST /api/v1/ai-gateway/moderations | Moderate content |
POST /api/v1/ai-gateway/rerank | Rerank documents against a query |
Each route's request and response are in the AI Inference reference.
These routes answer with the AI Gateway's own response body and status, as an OpenAI client expects them: they are not wrapped in the REST API's { data, meta } envelope. The X-Request-Id header is still set, so you can trace a request.
Managing models, keys and usage
Everything else about the gateway is a regular REST resource with the { data, meta } envelope:
| Area | Reference |
|---|---|
| Models: list, add, deploy, start, stop, resize, settings, metrics, logs | AI Models |
| Your providers' API keys | Provider Keys |
| Model routers (one name, several models) | Routers |
| Guardrails on a model | Guardrails |
| Usage, latency and cost analytics | AI Analytics |
| Fine-tuning jobs | Fine-Tuning |
| Data Forge (training data) | Data Forge |
Errors
Inference routes pass the gateway's error body and status through. A request the REST API itself refuses (no key, a key without the scope, a missing model) answers an RFC 9457 problem with type, title, status and detail:
| Status | Meaning |
|---|---|
401 | No API key, or not a valid one |
403 | The key lacks the scope, or you may not use the model |
404 | No such model (or one you cannot see) |
429 | Rate limited: retry after the Retry-After header |
502 | The AI Gateway could not be reached |
From inside the platform: STRONGLY_SERVICES
Apps, workspaces, jobs and workflow nodes get the platform's services in the STRONGLY_SERVICES environment variable. Its aigateway entry carries the gateway's in-cluster base_url and the models you chose for the workload (available_models). Call the gateway's OpenAI-compatible /v1/... routes at that address and name the model with its Strongly ID in X-Model-Id:
import json, os, requests
services = json.loads(os.environ["STRONGLY_SERVICES"])
gateway = services["services"]["aigateway"]
model = gateway["available_models"][0]
reply = requests.post(
f"{gateway['base_url']}/v1/chat/completions",
headers={"Content-Type": "application/json", "X-Model-Id": model["_id"]},
json={
"model": model["vendor_model_id"],
"messages": [{"role": "user", "content": "Hello"}],
},
)
print(reply.json()["choices"][0]["message"]["content"])
The platform authenticates these calls as the workload's owner, so they count toward that user's usage and budgets like their own requests. See Platform Services for every entry in STRONGLY_SERVICES and examples in Node.js.