AI Models
Manage AI models registered in the Strongly AI Gateway. Models can be third-party provider models (OpenAI, Anthropic, etc.) or self-hosted models deployed on your cluster.
AIModel Object
{
"_id": "model_abc123",
"name": "GPT-4 Production",
"type": "third-party",
"provider": "openai",
"vendorModelId": "gpt-4",
"modelType": "chat",
"status": "active",
"isActive": true,
"description": "GPT-4 for production workloads",
"capabilities": ["chat", "function-calling"],
"contextWindow": 128000,
"owner": "user_abc123",
"organizationId": "org_abc123",
"isShared": true,
"sharedWith": ["user_def456"],
"createdAt": "2025-01-15T10:00:00.000Z",
"updatedAt": "2025-02-01T14:30:00.000Z"
}
GET /api/v1/ai-gateway/models
List AI models
Returns a paginated list of AI models accessible to the current user.
Scope: ai-gateway:read
Parameters:
| Name | In | Type | Required | Description |
|---|---|---|---|---|
q | query | string | No | Search by model name |
type | query | string | No | Filter by type: third-party, self-hosted |
status | query | string | No | Filter by status: active, deploying, stopped, failed |
provider | query | string | No | Filter by provider: openai, anthropic, google, etc. |
modelType | query | string | No | Filter by model type: chat, completion, embedding |
limit | query | integer | No | Max results (default: 50, max: 200) |
cursor | query | string | No | meta.nextCursor of the previous page; omit for the first page |
sort | query | string | No | Sort field (default: -createdAt) |
Response: 200 OK (paginated)
{
"data": [
{
"_id": "model_abc123",
"name": "GPT-4 Production",
"type": "third-party",
"provider": "openai",
"vendorModelId": "gpt-4",
"modelType": "chat",
"status": "active",
"description": "GPT-4 for production workloads",
"capabilities": ["chat", "function-calling"],
"contextWindow": 128000,
"owner": "user_abc123",
"organizationId": "org_abc123",
"isShared": true,
"createdAt": "2025-01-15T10:00:00.000Z",
"updatedAt": "2025-02-01T14:30:00.000Z"
}
],
"meta": { "total": 12, "limit": 50, "nextCursor": null, "requestId": "req_abc123" }
}
GET /api/v1/ai-gateway/models/overview
Get model overview statistics
Returns aggregate counts of models by status and type.
Scope: ai-gateway:read
Response: 200 OK
{
"data": {
"total": 12,
"active": 8,
"deploying": 1,
"stopped": 2,
"failed": 1,
"thirdParty": 9,
"selfHosted": 3
},
"meta": { "requestId": "req_abc123" }
}
GET /api/v1/ai-gateway/models/certified
List curated third-party models
Returns the catalog of certified third-party models maintained by the AI Gateway (e.g. OpenAI, Anthropic, Mistral). Drives the mobile add-model picker so clients should never hardcode vendor/model lists.
Scope: ai-gateway:read
Parameters:
| Name | In | Type | Required | Description |
|---|---|---|---|---|
provider | query | string | No | Filter by provider (openai, anthropic, mistral, …) |
modelType | query | string | No | Filter by model type (chat, embedding, image, …) |
capability | query | string | No | Filter by capability flag (vision, audio, …) |
Response: 200 OK
{
"data": [
{
"modelId": "gpt-4o",
"displayName": "GPT-4o",
"modelType": "chat",
"provider": "openai",
"vendor": "OpenAI",
"capabilities": { "vision": true, "functionCalling": true },
"parameters": { "context_window": 128000 },
"apiEndpoint": "https://api.openai.com/v1"
}
],
"meta": { "requestId": "req_abc123" }
}
GET /api/v1/ai-gateway/models/providers
List third-party providers
Returns the distinct providers present in the certified models catalog, with a model count for each. Used by the mobile add-model vendor picker.
Scope: ai-gateway:read
Response: 200 OK
{
"data": [
{ "id": "openai", "label": "OpenAI", "count": 14 },
{ "id": "anthropic", "label": "Anthropic", "count": 6 }
],
"meta": { "requestId": "req_abc123" }
}
GET /api/v1/ai-gateway/models/prebuilt
List self-hosted prebuilt model templates
Returns the catalog of prebuilt model templates that can be deployed as self-hosted models on the cluster. Drives the self-hosted-deploy wizard's vendor/model picker.
Scope: ai-gateway:read
Parameters:
| Name | In | Type | Required | Description |
|---|---|---|---|---|
category | query | string | No | Filter by category (llm, image, audio, …) |
provider | query | string | No | Filter by provider (meta, google, microsoft, …) |
Response: 200 OK
{
"data": [
{
"id": "llama-3.1-8b-instruct",
"name": "Llama 3.1 8B Instruct",
"description": "Meta Llama 3.1 8B instruction-tuned chat model",
"category": "llm",
"provider": "meta",
"defaultPort": 8000,
"defaultResources": { "cpu": "2000m", "memory": "16Gi", "gpu": 1, "disk": "100Gi" },
"recommendedInstance": "g5.xlarge",
"modelSize": "8B",
"tags": ["chat", "instruct"]
}
],
"meta": { "requestId": "req_abc123" }
}
POST /api/v1/ai-gateway/models
Create a new AI model
Registers a new model in the AI Gateway. Third-party models become active immediately; self-hosted models require deployment.
Scope: ai-gateway:write
Request Body:
{
"name": "GPT-4 Production",
"type": "third-party",
"provider": "openai",
"vendorModelId": "gpt-4",
"modelType": "chat",
"description": "GPT-4 for production workloads",
"capabilities": ["chat", "function-calling"],
"maxTokens": 8192,
"contextWindow": 128000,
"config": {
"defaultTemperature": 0.7,
"rateLimit": 100
}
}
| Field | Type | Required | Description |
|---|---|---|---|
name | string | Yes | Human-readable model name |
type | string | Yes | Model type: third-party or self-hosted |
provider | string | Yes | Provider: openai, anthropic, google, huggingface, etc. |
vendorModelId | string | Yes | Provider's model identifier (e.g., gpt-6.1-sol, claude-opus-5-5) |
modelType | string | No | Capability type: chat, completion, embedding |
description | string | No | Model description |
capabilities | string[] | No | List of capabilities (e.g., chat, function-calling, vision) |
maxTokens | integer | No | Maximum output tokens |
contextWindow | integer | No | Maximum context window size in tokens |
config | object | No | Additional provider-specific configuration |
Response: 201 Created
{
"data": {
"_id": "model_abc123",
"name": "GPT-4 Production",
"type": "third-party",
"provider": "openai",
"vendorModelId": "gpt-4",
"modelType": "chat",
"status": "active",
"isActive": true,
"owner": "user_abc123",
"organizationId": "org_abc123",
"createdAt": "2025-02-07T10:00:00.000Z",
"updatedAt": "2025-02-07T10:00:00.000Z"
},
"meta": { "requestId": "req_abc123" }
}
GET /api/v1/ai-gateway/models/:id
Get an AI model
Returns the full details of a single model.
Scope: ai-gateway:read
Parameters:
| Name | In | Type | Required | Description |
|---|---|---|---|---|
id | path | string | Yes | Model ID |
Response: 200 OK
{
"data": {
"_id": "model_abc123",
"name": "GPT-4 Production",
"type": "third-party",
"provider": "openai",
"vendorModelId": "gpt-4",
"modelType": "chat",
"status": "active",
"isActive": true,
"description": "GPT-4 for production workloads",
"capabilities": ["chat", "function-calling"],
"contextWindow": 128000,
"owner": "user_abc123",
"organizationId": "org_abc123",
"isShared": true,
"sharedWith": ["user_def456"],
"createdAt": "2025-01-15T10:00:00.000Z",
"updatedAt": "2025-02-01T14:30:00.000Z"
},
"meta": { "requestId": "req_abc123" }
}
GET /api/v1/ai-gateway/models/:id/options
Get model-specific options
Returns runtime options exposed by the model (e.g. available voices for TTS, supported languages for STT, sampler choices). Proxies to the AI Gateway, which queries the model pod (self-hosted) or the provider API (third-party).
Scope: ai-gateway:read
Parameters:
| Name | In | Type | Required | Description |
|---|---|---|---|---|
id | path | string | Yes | Model ID |
Response: 200 OK
Shape is model-dependent. Typical examples:
{
"data": {
"voices": [
{ "id": "alloy", "label": "Alloy", "gender": "neutral" },
{ "id": "verse", "label": "Verse", "gender": "neutral" }
],
"languages": ["en", "es", "fr", "de", "ja"]
},
"meta": { "requestId": "req_abc123" }
}
Errors:
502 gateway-error-- AI Gateway unreachable or returned an error.500 config-error-- AI Gateway URL not configured on the server.
PATCH /api/v1/ai-gateway/models/:id
Update an AI model
Updates model properties: the same fields the model's Settings tab saves. Only provided fields are changed.
For a deployed self-hosted model, replicas, the scheduling fields and the scaling fields change the running model without a redeploy. They are checked together against the model as it is now, and an invalid result is refused with 400 validation-error and nothing changes:
- An on-demand model needs an idle window (
autoShutdownMinutes, an integer >= 1). - A scheduled model needs
scheduleRuleswith at least one start and one stop time. scalingModeis exactlyfixedorauto.autoneeds an on-demand model and all three ofminReplicas,maxReplicasandtargetConcurrency, withminReplicasno greater thanmaxReplicas.
To change CPU, memory, disk, GPU or spot capacity, use PATCH /api/v1/ai-gateway/models/:id/resources. Sharing (sharedWith, usedBy, isShared) is set through the model's permissions; an update that includes them is refused with 400.
Scope: ai-gateway:write
Parameters:
| Name | In | Type | Required | Description |
|---|---|---|---|---|
id | path | string | Yes | Model ID |
Request Body:
{
"name": "GPT-4 Production (Updated)",
"description": "Updated description",
"maxTokens": 4096,
"config": {
"defaultTemperature": 0.5
}
}
Switching a self-hosted model to on-demand with demand scaling:
{
"schedulingMode": "on_demand",
"autoShutdownMinutes": 15,
"scalingMode": "auto",
"minReplicas": 1,
"maxReplicas": 3,
"targetConcurrency": 4
}
| Field | Type | Required | Description |
|---|---|---|---|
name | string | No | Updated model name |
description | string | No | Updated description |
capabilities | string[] | No | Updated capabilities list |
maxTokens | integer | No | Updated max output tokens |
contextWindow | integer | No | Updated context window size |
config | object | No | Updated configuration (merged with existing) |
cacheConfig | object | No | Caching: { semantic_cache_enabled, semantic_cache_threshold (0-1), semantic_cache_ttl (seconds) } |
replicas | integer | No | Self-hosted: pods while the model runs (>= 1). On-demand models wake to this many |
schedulingMode | string | No | Self-hosted: always_on, on_demand or scheduled |
autoShutdownMinutes | integer | No | Self-hosted, on-demand: idle minutes before the model scales to zero (>= 1) |
scheduleRules | object[] | No | Self-hosted, scheduled: start/stop times, each { id, action: "start" | "stop", dayOfWeek: [0-6, 0 = Sunday], time: "HH:MM", timezone } (IANA timezone) |
scalingMode | string | No | Self-hosted: fixed (wake to replicas) or auto (replicas follow live load). Any other value is refused |
minReplicas | integer | No | auto only: fewest replicas while awake (>= 1) |
maxReplicas | integer | No | auto only: most replicas (>= minReplicas) |
targetConcurrency | integer | No | auto only: in-flight requests one replica serves before another is added (>= 1) |
Response: 200 OK: the model, as GET /ai-gateway/models/:id shows it.
Errors:
400 validation-error-- A field is invalid or the resulting scheduling/scaling combination is not coherent; the message names the field and the value received.403 forbidden-- You can update only your own models (admins can update any).404 not-found-- No such model.
PATCH /api/v1/ai-gateway/models/:id/resources
Resize a deployed self-hosted model in place
Changes the CPU, memory, disk and/or GPU count of a running self-hosted model, and/or moves it on or off spot capacity: the same save as the Environment tab's resource fields and spot setting. The model restarts at exactly the size and capacity you send (what you request is what it gets); it is not deleted or redeployed, and its endpoint stays the same. A bigger size costs more, so the model's budget is checked first; a refused resize changes nothing.
Scope: ai-gateway:write
Parameters:
| Name | In | Type | Required | Description |
|---|---|---|---|---|
id | path | string | Yes | Model ID |
cpu | body | string | No | CPU, e.g. "2" or "2000m" |
memory | body | string | No | Memory, e.g. "8Gi" |
disk | body | string | No | Disk, e.g. "100Gi" |
gpu | body | integer | No | GPU count |
useSpot | body | boolean | No | Run on spot capacity (cheaper, can be reclaimed: the model restarts on a new node) |
spotFallback | body | boolean | No | With useSpot: fall back to on-demand when no spot capacity is available (default true); false waits for spot |
Send at least one of cpu, memory, disk, gpu or useSpot.
Request example:
{
"cpu": "4",
"memory": "16Gi"
}
Response: 200 OK: the model, as GET /ai-gateway/models/:id shows it, at its new size.
Errors:
400 validation-error-- No field sent, a value is invalid, or the model is not a deployed self-hosted model.402 payment-required-- A budget blocks the new size; the message is the budget's reason.403 forbidden-- You cannot edit this model.404 not-found-- No such model.
DELETE /api/v1/ai-gateway/models/:id
Delete an AI model
Permanently removes a model from the AI Gateway. Self-hosted models are undeployed first.
Scope: ai-gateway:write
Parameters:
| Name | In | Type | Required | Description |
|---|---|---|---|---|
id | path | string | Yes | Model ID |
Response: 204 No Content
DELETE /api/v1/ai-gateway/models/:id/cache
Clear the model's semantic cache
Removes all semantic-cache entries for this model. Future requests will miss the cache and generate fresh responses. Useful after changing the system prompt, model parameters, or cache configuration.
Scope: ai-gateway:write
Parameters:
| Name | In | Type | Required | Description |
|---|---|---|---|---|
id | path | string | Yes | Model ID |
Response: 204 No Content
A model you may not change, or no model with that id, is 404 not-found.
POST /api/v1/ai-gateway/models/:id/deploy
Deploy a model
Deploys a self-hosted model to the cluster (for a third-party model, activates it). On a self-hosted model that is already deployed, it applies any scheduling, scaling, replica, instance, spot or environment changes you send and starts the model.
How a self-hosted model runs:
- Always-on (default):
replicaspods run continuously. - On-demand: set
autoShutdownMinutes>= 1. The model scales to zero after that many idle minutes and wakes on the next request (the first request after idle waits for the model to start). - Scheduled: send
scheduleRules; the model starts and stops at those times and runsreplicaspods in between. - Demand scaling (
scaling_mode: "auto", on-demand only): while awake, replicas follow live load betweenminReplicasandmaxReplicas, adding a replica for everytargetConcurrencyrequests in flight.scaling_mode: "fixed"(default) always wakes toreplicas.
Every field is checked before anything is deployed or changed; an invalid value or combination is refused with 400 validation-error. First deploys also pass the same governance requirements and budget check as a deploy from the UI.
Scope: ai-gateway:write
Parameters:
| Name | In | Type | Required | Description |
|---|---|---|---|---|
id | path | string | Yes | Model ID |
instanceType | body | string | No | EC2 instance type to pin (e.g. g5.xlarge). Otherwise the model's recommended instance is used; required when the model has none |
replicas | body | integer | No | Pods while the model runs, >= 1. Default 1 |
autoShutdownMinutes | body | integer | No | >= 1 makes the model on-demand with this idle window. 0 or omitted: always-on (on a deployed model, 0 switches it to always-on) |
scheduleRules | body | object[] | No | Scheduled mode: start/stop times, each { id, action: "start" | "stop", dayOfWeek: [0-6, 0 = Sunday], time: "HH:MM", timezone }, with at least one start and one stop. Not combinable with autoShutdownMinutes |
scalingMode | body | string | No | Exactly fixed (default) or auto. auto needs autoShutdownMinutes >= 1 and all three fields below |
minReplicas | body | integer | No | auto only: fewest replicas while awake, >= 1 |
maxReplicas | body | integer | No | auto only: most replicas, >= minReplicas |
targetConcurrency | body | integer | No | auto only: in-flight requests one replica serves before another is added, >= 1 |
envVars | body | object | No | Environment variables for the model container |
useSpot | body | boolean | No | Schedule on spot (interruptible) capacity -- up to 70% cheaper. Default false. See Spot Instance Support |
spotFallback | body | boolean | No | With useSpot, request spot as a scheduling preference and fall back to on-demand when spot capacity is unavailable. Default true; set false to pin strictly to spot |
Request example:
{
"instanceType": "g5.xlarge",
"replicas": 1,
"autoShutdownMinutes": 10,
"scalingMode": "auto",
"minReplicas": 1,
"maxReplicas": 4,
"targetConcurrency": 3,
"useSpot": true,
"spotFallback": true
}
Response: 200 OK. A first self-hosted deploy returns the model with status: "deploying" and the stored settings (schedulingMode, autoShutdownMinutes, replicas, scalingMode, bounds). Deploying an already deployed model returns the start result.
{
"data": {
"_id": "model_abc123",
"status": "deploying",
"schedulingMode": "on_demand",
"autoShutdownMinutes": 10,
"replicas": 1,
"scalingMode": "auto",
"minReplicas": 1,
"maxReplicas": 4,
"targetConcurrency": 3
},
"meta": { "requestId": "req_abc123" }
}
Errors:
400 validation-error-- A field is invalid, e.g.scaling_mode: "demand",scaling_mode: "auto"on an always-on model or without its three bounds, bounds sent withoutauto,replicas: 0, or an invertedminReplicas/maxReplicas. The message names the field and the value received.402 payment-required-- A budget blocks the deploy.403 governance-blocked-- The model's governance requirements are not yet met; the message says which.404 not-found-- No such model, or you cannot edit it.
POST /api/v1/ai-gateway/models/:id/start
Start a stopped model
Starts a previously stopped self-hosted model.
Scope: ai-gateway:write
Parameters:
| Name | In | Type | Required | Description |
|---|---|---|---|---|
id | path | string | Yes | Model ID |
Response: 200 OK: the model, as GET /ai-gateway/models/:id shows it, its status now (starting, then running; failed with the reason when its pod does not become ready). A third-party model is activated (status active).
POST /api/v1/ai-gateway/models/:id/stop
Stop a running model
Stops a running self-hosted model, freeing cluster resources. A third-party model is deactivated instead.
Scope: ai-gateway:write
Parameters:
| Name | In | Type | Required | Description |
|---|---|---|---|---|
id | path | string | Yes | Model ID |
Response: 200 OK: the model, as GET /ai-gateway/models/:id shows it, status stopped (a third-party model: inactive).
Errors: if the model cannot be stopped, the response carries the AI Gateway's reason and status (for example 500 when the model's deployment cannot be scaled down), and 502 when the AI Gateway cannot be reached. A stop that failed does not mark the model stopped.
GET /api/v1/ai-gateway/models/:id/status
Get model deployment status
Returns the current deployment status of a model.
Scope: ai-gateway:read
Parameters:
| Name | In | Type | Required | Description |
|---|---|---|---|---|
id | path | string | Yes | Model ID |
Response: 200 OK
For a self-hosted model:
{
"data": {
"deploymentId": "llama-3-1-8b-a1b2c3",
"name": "Llama 3.1 8B",
"status": "running",
"error": null,
"resources": { "cpu": "4", "memory": "16Gi", "gpu": 1, "disk": "100Gi" },
"endpoints": { "public": null, "private": "http://<internal-host>:8000" },
"createdAt": "2025-02-07T10:00:00.000Z",
"lastRequestAt": "2025-02-07T11:42:10.000Z",
"autoShutdownAt": null
},
"meta": { "requestId": "req_abc123" }
}
For a third-party model:
{
"data": {
"status": "active",
"type": "third-party",
"name": "GPT-4 Production",
"provider": "openai",
"isActive": true,
"updatedAt": "2025-02-01T14:30:00.000Z"
},
"meta": { "requestId": "req_abc123" }
}
GET /api/v1/ai-gateway/models/:id/metrics
Get model metrics
For a self-hosted model: its running replicas measured live (CPU, memory, GPU, the model cache on disk, network, replicas). For any other model: the requests the AI Gateway served for it over a window.
Scope: ai-gateway:read
Parameters:
| Name | In | Type | Required | Description |
|---|---|---|---|---|
id | path | string | Yes | Model ID |
timeRange | query | string | No | Third-party models: 1h, 6h, 24h or 7d (default 24h) |
Response: 200 OK
For a self-hosted model:
{
"data": {
"deploymentId": "llama-3-1-8b-a1b2c3",
"cpu": { "usageMillicores": 812.4, "limitMillicores": 4000, "percent": 20.3 },
"memory": { "usageBytes": 9126805504, "limitBytes": 17179869184, "percent": 53.1 },
"gpu": {
"count": 1,
"utilizationPercent": 41.0,
"memoryUsedBytes": 20401094656,
"memoryTotalBytes": 24146608128,
"memoryPercent": 84.5,
"devices": [ /* one entry per GPU */ ]
},
"disk": { "path": "/model_cache", "usedBytes": 16106127360, "limitBytes": 107374182400, "percent": 15.0 },
"network": { "receiveBytesPerSecond": 2048.0, "transmitBytesPerSecond": 8192.5, "totalBytesPerSecond": 10240.5, "windowSeconds": 5.0 },
"replicas": { "desired": 1, "ready": 1, "measured": 1 },
"measuredAt": "2025-02-07T12:00:00.000000+00:00"
},
"meta": { "requestId": "req_abc123" }
}
gpu is null on a model without GPUs. A model with nothing running returns 409.
For a third-party model:
{
"data": {
"modelId": "model_abc123",
"timeRange": "24h",
"windowStart": "2025-02-06T12:00:00Z",
"coveredSince": "2025-02-06T12:00:00Z",
"windowEnd": "2025-02-07T12:00:00Z",
"requests": 15420,
"failed": 31,
"requestsPerMinute": 10.71,
"avgResponseTimeMs": 450.2,
"totalTokens": 2500000
},
"meta": { "requestId": "req_abc123" }
}
avgResponseTimeMs and totalTokens are null when no request in the window recorded them. When older records have been archived, the figures cover coveredSince to now. The figures count the requests you may count: a platform administrator every request, an organization owner or admin on a multi-tenant platform their organization's members', and anyone else their own (see Monitoring).
GET /api/v1/ai-gateway/models/:id/logs
Get model logs
Returns container logs for a self-hosted model deployment.
Scope: ai-gateway:read
Parameters:
| Name | In | Type | Required | Description |
|---|---|---|---|---|
id | path | string | Yes | Model ID |
lines | query | integer | No | Number of log lines to return (default: 100) |
since | query | string | No | Return logs since this time (ISO 8601 or duration like 1h, 30m) |
container | query | string | No | Container name (for multi-container pods) |
Response: 200 OK
For a self-hosted model, data is the list of log entries:
{
"data": [
{ "timestamp": "2025-02-07T10:00:00Z", "level": "info", "message": "Model loaded successfully" },
{ "timestamp": "2025-02-07T10:00:01Z", "level": "info", "message": "Ready to serve requests on port 8000" }
],
"meta": { "requestId": "req_abc123" }
}
A third-party model has no deployment logs; data is { "logs": [], "message": "Third-party models do not generate deployment logs" }.
Sharing
Who can reach a AI model is the platform's one sharing shape: its owner, its members
(role editor can use and change it, user can only use it) and its visibility
(public: every user can find and use it; in a multi-tenant deployment, every user of
its organization). Changing it stays with its owner and editors.
| Method | Path | Does | Scope |
|---|---|---|---|
| GET | /api/v1/ai-gateway/models/:id/permissions | Owner, members (userId, role) and visibility | ai-gateway:read |
| POST | /api/v1/ai-gateway/models/:id/permissions/members | Share with a user: { "userId", "role": "editor" | "user" } | ai-gateway:write |
| DELETE | /api/v1/ai-gateway/models/:aiModelId/permissions/members/:userId | Stop sharing with a user | ai-gateway:write |
| PATCH | /api/v1/ai-gateway/models/:id/permissions | { "visibility": "public" | "private" } | ai-gateway:write |
Each, except a member's removal (204), answers the permissions as they are now:
{
"data": {
"resourceId": "<AI model id>",
"owner": "<user id>",
"members": [{ "userId": "<user id>", "role": "user" }],
"visibility": "private"
},
"meta": { "requestId": "req_abc123" }
}