Skip to main content

AI Models

Manage AI models registered in the Strongly AI Gateway. Models can be third-party provider models (OpenAI, Anthropic, etc.) or self-hosted models deployed on your cluster.


AIModel Object​

{
"_id": "model_abc123",
"name": "GPT-4 Production",
"type": "third-party",
"provider": "openai",
"vendorModelId": "gpt-4",
"modelType": "chat",
"status": "active",
"isActive": true,
"description": "GPT-4 for production workloads",
"capabilities": ["chat", "function-calling"],
"contextWindow": 128000,
"owner": "user_abc123",
"organizationId": "org_abc123",
"isShared": true,
"sharedWith": ["user_def456"],
"createdAt": "2025-01-15T10:00:00.000Z",
"updatedAt": "2025-02-01T14:30:00.000Z"
}

GET /api/v1/ai-gateway/models​

List AI models

Returns a paginated list of AI models accessible to the current user.

Scope: ai-gateway:read

Parameters:

NameInTypeRequiredDescription
qquerystringNoSearch by model name
typequerystringNoFilter by type: third-party, self-hosted
statusquerystringNoFilter by status: active, deploying, stopped, failed
providerquerystringNoFilter by provider: openai, anthropic, google, etc.
modelTypequerystringNoFilter by model type: chat, completion, embedding
limitqueryintegerNoMax results (default: 50, max: 200)
cursorquerystringNometa.nextCursor of the previous page; omit for the first page
sortquerystringNoSort field (default: -createdAt)

Response: 200 OK (paginated)

{
"data": [
{
"_id": "model_abc123",
"name": "GPT-4 Production",
"type": "third-party",
"provider": "openai",
"vendorModelId": "gpt-4",
"modelType": "chat",
"status": "active",
"description": "GPT-4 for production workloads",
"capabilities": ["chat", "function-calling"],
"contextWindow": 128000,
"owner": "user_abc123",
"organizationId": "org_abc123",
"isShared": true,
"createdAt": "2025-01-15T10:00:00.000Z",
"updatedAt": "2025-02-01T14:30:00.000Z"
}
],
"meta": { "total": 12, "limit": 50, "nextCursor": null, "requestId": "req_abc123" }
}

GET /api/v1/ai-gateway/models/overview​

Get model overview statistics

Returns aggregate counts of models by status and type.

Scope: ai-gateway:read

Response: 200 OK

{
"data": {
"total": 12,
"active": 8,
"deploying": 1,
"stopped": 2,
"failed": 1,
"thirdParty": 9,
"selfHosted": 3
},
"meta": { "requestId": "req_abc123" }
}

GET /api/v1/ai-gateway/models/certified​

List curated third-party models

Returns the catalog of certified third-party models maintained by the AI Gateway (e.g. OpenAI, Anthropic, Mistral). Drives the mobile add-model picker so clients should never hardcode vendor/model lists.

Scope: ai-gateway:read

Parameters:

NameInTypeRequiredDescription
providerquerystringNoFilter by provider (openai, anthropic, mistral, …)
modelTypequerystringNoFilter by model type (chat, embedding, image, …)
capabilityquerystringNoFilter by capability flag (vision, audio, …)

Response: 200 OK

{
"data": [
{
"modelId": "gpt-4o",
"displayName": "GPT-4o",
"modelType": "chat",
"provider": "openai",
"vendor": "OpenAI",
"capabilities": { "vision": true, "functionCalling": true },
"parameters": { "context_window": 128000 },
"apiEndpoint": "https://api.openai.com/v1"
}
],
"meta": { "requestId": "req_abc123" }
}

GET /api/v1/ai-gateway/models/providers​

List third-party providers

Returns the distinct providers present in the certified models catalog, with a model count for each. Used by the mobile add-model vendor picker.

Scope: ai-gateway:read

Response: 200 OK

{
"data": [
{ "id": "openai", "label": "OpenAI", "count": 14 },
{ "id": "anthropic", "label": "Anthropic", "count": 6 }
],
"meta": { "requestId": "req_abc123" }
}

GET /api/v1/ai-gateway/models/prebuilt​

List self-hosted prebuilt model templates

Returns the catalog of prebuilt model templates that can be deployed as self-hosted models on the cluster. Drives the self-hosted-deploy wizard's vendor/model picker.

Scope: ai-gateway:read

Parameters:

NameInTypeRequiredDescription
categoryquerystringNoFilter by category (llm, image, audio, …)
providerquerystringNoFilter by provider (meta, google, microsoft, …)

Response: 200 OK

{
"data": [
{
"id": "llama-3.1-8b-instruct",
"name": "Llama 3.1 8B Instruct",
"description": "Meta Llama 3.1 8B instruction-tuned chat model",
"category": "llm",
"provider": "meta",
"defaultPort": 8000,
"defaultResources": { "cpu": "2000m", "memory": "16Gi", "gpu": 1, "disk": "100Gi" },
"recommendedInstance": "g5.xlarge",
"modelSize": "8B",
"tags": ["chat", "instruct"]
}
],
"meta": { "requestId": "req_abc123" }
}

POST /api/v1/ai-gateway/models​

Create a new AI model

Registers a new model in the AI Gateway. Third-party models become active immediately; self-hosted models require deployment.

Scope: ai-gateway:write

Request Body:

{
"name": "GPT-4 Production",
"type": "third-party",
"provider": "openai",
"vendorModelId": "gpt-4",
"modelType": "chat",
"description": "GPT-4 for production workloads",
"capabilities": ["chat", "function-calling"],
"maxTokens": 8192,
"contextWindow": 128000,
"config": {
"defaultTemperature": 0.7,
"rateLimit": 100
}
}
FieldTypeRequiredDescription
namestringYesHuman-readable model name
typestringYesModel type: third-party or self-hosted
providerstringYesProvider: openai, anthropic, google, huggingface, etc.
vendorModelIdstringYesProvider's model identifier (e.g., gpt-6.1-sol, claude-opus-5-5)
modelTypestringNoCapability type: chat, completion, embedding
descriptionstringNoModel description
capabilitiesstring[]NoList of capabilities (e.g., chat, function-calling, vision)
maxTokensintegerNoMaximum output tokens
contextWindowintegerNoMaximum context window size in tokens
configobjectNoAdditional provider-specific configuration

Response: 201 Created

{
"data": {
"_id": "model_abc123",
"name": "GPT-4 Production",
"type": "third-party",
"provider": "openai",
"vendorModelId": "gpt-4",
"modelType": "chat",
"status": "active",
"isActive": true,
"owner": "user_abc123",
"organizationId": "org_abc123",
"createdAt": "2025-02-07T10:00:00.000Z",
"updatedAt": "2025-02-07T10:00:00.000Z"
},
"meta": { "requestId": "req_abc123" }
}

GET /api/v1/ai-gateway/models/:id​

Get an AI model

Returns the full details of a single model.

Scope: ai-gateway:read

Parameters:

NameInTypeRequiredDescription
idpathstringYesModel ID

Response: 200 OK

{
"data": {
"_id": "model_abc123",
"name": "GPT-4 Production",
"type": "third-party",
"provider": "openai",
"vendorModelId": "gpt-4",
"modelType": "chat",
"status": "active",
"isActive": true,
"description": "GPT-4 for production workloads",
"capabilities": ["chat", "function-calling"],
"contextWindow": 128000,
"owner": "user_abc123",
"organizationId": "org_abc123",
"isShared": true,
"sharedWith": ["user_def456"],
"createdAt": "2025-01-15T10:00:00.000Z",
"updatedAt": "2025-02-01T14:30:00.000Z"
},
"meta": { "requestId": "req_abc123" }
}

GET /api/v1/ai-gateway/models/:id/options​

Get model-specific options

Returns runtime options exposed by the model (e.g. available voices for TTS, supported languages for STT, sampler choices). Proxies to the AI Gateway, which queries the model pod (self-hosted) or the provider API (third-party).

Scope: ai-gateway:read

Parameters:

NameInTypeRequiredDescription
idpathstringYesModel ID

Response: 200 OK

Shape is model-dependent. Typical examples:

{
"data": {
"voices": [
{ "id": "alloy", "label": "Alloy", "gender": "neutral" },
{ "id": "verse", "label": "Verse", "gender": "neutral" }
],
"languages": ["en", "es", "fr", "de", "ja"]
},
"meta": { "requestId": "req_abc123" }
}

Errors:

  • 502 gateway-error -- AI Gateway unreachable or returned an error.
  • 500 config-error -- AI Gateway URL not configured on the server.

PATCH /api/v1/ai-gateway/models/:id​

Update an AI model

Updates model properties: the same fields the model's Settings tab saves. Only provided fields are changed.

For a deployed self-hosted model, replicas, the scheduling fields and the scaling fields change the running model without a redeploy. They are checked together against the model as it is now, and an invalid result is refused with 400 validation-error and nothing changes:

  • An on-demand model needs an idle window (autoShutdownMinutes, an integer >= 1).
  • A scheduled model needs scheduleRules with at least one start and one stop time.
  • scalingMode is exactly fixed or auto. auto needs an on-demand model and all three of minReplicas, maxReplicas and targetConcurrency, with minReplicas no greater than maxReplicas.

To change CPU, memory, disk, GPU or spot capacity, use PATCH /api/v1/ai-gateway/models/:id/resources. Sharing (sharedWith, usedBy, isShared) is set through the model's permissions; an update that includes them is refused with 400.

Scope: ai-gateway:write

Parameters:

NameInTypeRequiredDescription
idpathstringYesModel ID

Request Body:

{
"name": "GPT-4 Production (Updated)",
"description": "Updated description",
"maxTokens": 4096,
"config": {
"defaultTemperature": 0.5
}
}

Switching a self-hosted model to on-demand with demand scaling:

{
"schedulingMode": "on_demand",
"autoShutdownMinutes": 15,
"scalingMode": "auto",
"minReplicas": 1,
"maxReplicas": 3,
"targetConcurrency": 4
}
FieldTypeRequiredDescription
namestringNoUpdated model name
descriptionstringNoUpdated description
capabilitiesstring[]NoUpdated capabilities list
maxTokensintegerNoUpdated max output tokens
contextWindowintegerNoUpdated context window size
configobjectNoUpdated configuration (merged with existing)
cacheConfigobjectNoCaching: { semantic_cache_enabled, semantic_cache_threshold (0-1), semantic_cache_ttl (seconds) }
replicasintegerNoSelf-hosted: pods while the model runs (>= 1). On-demand models wake to this many
schedulingModestringNoSelf-hosted: always_on, on_demand or scheduled
autoShutdownMinutesintegerNoSelf-hosted, on-demand: idle minutes before the model scales to zero (>= 1)
scheduleRulesobject[]NoSelf-hosted, scheduled: start/stop times, each { id, action: "start" | "stop", dayOfWeek: [0-6, 0 = Sunday], time: "HH:MM", timezone } (IANA timezone)
scalingModestringNoSelf-hosted: fixed (wake to replicas) or auto (replicas follow live load). Any other value is refused
minReplicasintegerNoauto only: fewest replicas while awake (>= 1)
maxReplicasintegerNoauto only: most replicas (>= minReplicas)
targetConcurrencyintegerNoauto only: in-flight requests one replica serves before another is added (>= 1)

Response: 200 OK: the model, as GET /ai-gateway/models/:id shows it.

Errors:

  • 400 validation-error -- A field is invalid or the resulting scheduling/scaling combination is not coherent; the message names the field and the value received.
  • 403 forbidden -- You can update only your own models (admins can update any).
  • 404 not-found -- No such model.

PATCH /api/v1/ai-gateway/models/:id/resources​

Resize a deployed self-hosted model in place

Changes the CPU, memory, disk and/or GPU count of a running self-hosted model, and/or moves it on or off spot capacity: the same save as the Environment tab's resource fields and spot setting. The model restarts at exactly the size and capacity you send (what you request is what it gets); it is not deleted or redeployed, and its endpoint stays the same. A bigger size costs more, so the model's budget is checked first; a refused resize changes nothing.

Scope: ai-gateway:write

Parameters:

NameInTypeRequiredDescription
idpathstringYesModel ID
cpubodystringNoCPU, e.g. "2" or "2000m"
memorybodystringNoMemory, e.g. "8Gi"
diskbodystringNoDisk, e.g. "100Gi"
gpubodyintegerNoGPU count
useSpotbodybooleanNoRun on spot capacity (cheaper, can be reclaimed: the model restarts on a new node)
spotFallbackbodybooleanNoWith useSpot: fall back to on-demand when no spot capacity is available (default true); false waits for spot

Send at least one of cpu, memory, disk, gpu or useSpot.

Request example:

{
"cpu": "4",
"memory": "16Gi"
}

Response: 200 OK: the model, as GET /ai-gateway/models/:id shows it, at its new size.

Errors:

  • 400 validation-error -- No field sent, a value is invalid, or the model is not a deployed self-hosted model.
  • 402 payment-required -- A budget blocks the new size; the message is the budget's reason.
  • 403 forbidden -- You cannot edit this model.
  • 404 not-found -- No such model.

DELETE /api/v1/ai-gateway/models/:id​

Delete an AI model

Permanently removes a model from the AI Gateway. Self-hosted models are undeployed first.

Scope: ai-gateway:write

Parameters:

NameInTypeRequiredDescription
idpathstringYesModel ID

Response: 204 No Content


DELETE /api/v1/ai-gateway/models/:id/cache​

Clear the model's semantic cache

Removes all semantic-cache entries for this model. Future requests will miss the cache and generate fresh responses. Useful after changing the system prompt, model parameters, or cache configuration.

Scope: ai-gateway:write

Parameters:

NameInTypeRequiredDescription
idpathstringYesModel ID

Response: 204 No Content

A model you may not change, or no model with that id, is 404 not-found.


POST /api/v1/ai-gateway/models/:id/deploy​

Deploy a model

Deploys a self-hosted model to the cluster (for a third-party model, activates it). On a self-hosted model that is already deployed, it applies any scheduling, scaling, replica, instance, spot or environment changes you send and starts the model.

How a self-hosted model runs:

  • Always-on (default): replicas pods run continuously.
  • On-demand: set autoShutdownMinutes >= 1. The model scales to zero after that many idle minutes and wakes on the next request (the first request after idle waits for the model to start).
  • Scheduled: send scheduleRules; the model starts and stops at those times and runs replicas pods in between.
  • Demand scaling (scaling_mode: "auto", on-demand only): while awake, replicas follow live load between minReplicas and maxReplicas, adding a replica for every targetConcurrency requests in flight. scaling_mode: "fixed" (default) always wakes to replicas.

Every field is checked before anything is deployed or changed; an invalid value or combination is refused with 400 validation-error. First deploys also pass the same governance requirements and budget check as a deploy from the UI.

Scope: ai-gateway:write

Parameters:

NameInTypeRequiredDescription
idpathstringYesModel ID
instanceTypebodystringNoEC2 instance type to pin (e.g. g5.xlarge). Otherwise the model's recommended instance is used; required when the model has none
replicasbodyintegerNoPods while the model runs, >= 1. Default 1
autoShutdownMinutesbodyintegerNo>= 1 makes the model on-demand with this idle window. 0 or omitted: always-on (on a deployed model, 0 switches it to always-on)
scheduleRulesbodyobject[]NoScheduled mode: start/stop times, each { id, action: "start" | "stop", dayOfWeek: [0-6, 0 = Sunday], time: "HH:MM", timezone }, with at least one start and one stop. Not combinable with autoShutdownMinutes
scalingModebodystringNoExactly fixed (default) or auto. auto needs autoShutdownMinutes >= 1 and all three fields below
minReplicasbodyintegerNoauto only: fewest replicas while awake, >= 1
maxReplicasbodyintegerNoauto only: most replicas, >= minReplicas
targetConcurrencybodyintegerNoauto only: in-flight requests one replica serves before another is added, >= 1
envVarsbodyobjectNoEnvironment variables for the model container
useSpotbodybooleanNoSchedule on spot (interruptible) capacity -- up to 70% cheaper. Default false. See Spot Instance Support
spotFallbackbodybooleanNoWith useSpot, request spot as a scheduling preference and fall back to on-demand when spot capacity is unavailable. Default true; set false to pin strictly to spot

Request example:

{
"instanceType": "g5.xlarge",
"replicas": 1,
"autoShutdownMinutes": 10,
"scalingMode": "auto",
"minReplicas": 1,
"maxReplicas": 4,
"targetConcurrency": 3,
"useSpot": true,
"spotFallback": true
}

Response: 200 OK. A first self-hosted deploy returns the model with status: "deploying" and the stored settings (schedulingMode, autoShutdownMinutes, replicas, scalingMode, bounds). Deploying an already deployed model returns the start result.

{
"data": {
"_id": "model_abc123",
"status": "deploying",
"schedulingMode": "on_demand",
"autoShutdownMinutes": 10,
"replicas": 1,
"scalingMode": "auto",
"minReplicas": 1,
"maxReplicas": 4,
"targetConcurrency": 3
},
"meta": { "requestId": "req_abc123" }
}

Errors:

  • 400 validation-error -- A field is invalid, e.g. scaling_mode: "demand", scaling_mode: "auto" on an always-on model or without its three bounds, bounds sent without auto, replicas: 0, or an inverted minReplicas/maxReplicas. The message names the field and the value received.
  • 402 payment-required -- A budget blocks the deploy.
  • 403 governance-blocked -- The model's governance requirements are not yet met; the message says which.
  • 404 not-found -- No such model, or you cannot edit it.

POST /api/v1/ai-gateway/models/:id/start​

Start a stopped model

Starts a previously stopped self-hosted model.

Scope: ai-gateway:write

Parameters:

NameInTypeRequiredDescription
idpathstringYesModel ID

Response: 200 OK: the model, as GET /ai-gateway/models/:id shows it, its status now (starting, then running; failed with the reason when its pod does not become ready). A third-party model is activated (status active).


POST /api/v1/ai-gateway/models/:id/stop​

Stop a running model

Stops a running self-hosted model, freeing cluster resources. A third-party model is deactivated instead.

Scope: ai-gateway:write

Parameters:

NameInTypeRequiredDescription
idpathstringYesModel ID

Response: 200 OK: the model, as GET /ai-gateway/models/:id shows it, status stopped (a third-party model: inactive).

Errors: if the model cannot be stopped, the response carries the AI Gateway's reason and status (for example 500 when the model's deployment cannot be scaled down), and 502 when the AI Gateway cannot be reached. A stop that failed does not mark the model stopped.


GET /api/v1/ai-gateway/models/:id/status​

Get model deployment status

Returns the current deployment status of a model.

Scope: ai-gateway:read

Parameters:

NameInTypeRequiredDescription
idpathstringYesModel ID

Response: 200 OK

For a self-hosted model:

{
"data": {
"deploymentId": "llama-3-1-8b-a1b2c3",
"name": "Llama 3.1 8B",
"status": "running",
"error": null,
"resources": { "cpu": "4", "memory": "16Gi", "gpu": 1, "disk": "100Gi" },
"endpoints": { "public": null, "private": "http://<internal-host>:8000" },
"createdAt": "2025-02-07T10:00:00.000Z",
"lastRequestAt": "2025-02-07T11:42:10.000Z",
"autoShutdownAt": null
},
"meta": { "requestId": "req_abc123" }
}

For a third-party model:

{
"data": {
"status": "active",
"type": "third-party",
"name": "GPT-4 Production",
"provider": "openai",
"isActive": true,
"updatedAt": "2025-02-01T14:30:00.000Z"
},
"meta": { "requestId": "req_abc123" }
}

GET /api/v1/ai-gateway/models/:id/metrics​

Get model metrics

For a self-hosted model: its running replicas measured live (CPU, memory, GPU, the model cache on disk, network, replicas). For any other model: the requests the AI Gateway served for it over a window.

Scope: ai-gateway:read

Parameters:

NameInTypeRequiredDescription
idpathstringYesModel ID
timeRangequerystringNoThird-party models: 1h, 6h, 24h or 7d (default 24h)

Response: 200 OK

For a self-hosted model:

{
"data": {
"deploymentId": "llama-3-1-8b-a1b2c3",
"cpu": { "usageMillicores": 812.4, "limitMillicores": 4000, "percent": 20.3 },
"memory": { "usageBytes": 9126805504, "limitBytes": 17179869184, "percent": 53.1 },
"gpu": {
"count": 1,
"utilizationPercent": 41.0,
"memoryUsedBytes": 20401094656,
"memoryTotalBytes": 24146608128,
"memoryPercent": 84.5,
"devices": [ /* one entry per GPU */ ]
},
"disk": { "path": "/model_cache", "usedBytes": 16106127360, "limitBytes": 107374182400, "percent": 15.0 },
"network": { "receiveBytesPerSecond": 2048.0, "transmitBytesPerSecond": 8192.5, "totalBytesPerSecond": 10240.5, "windowSeconds": 5.0 },
"replicas": { "desired": 1, "ready": 1, "measured": 1 },
"measuredAt": "2025-02-07T12:00:00.000000+00:00"
},
"meta": { "requestId": "req_abc123" }
}

gpu is null on a model without GPUs. A model with nothing running returns 409.

For a third-party model:

{
"data": {
"modelId": "model_abc123",
"timeRange": "24h",
"windowStart": "2025-02-06T12:00:00Z",
"coveredSince": "2025-02-06T12:00:00Z",
"windowEnd": "2025-02-07T12:00:00Z",
"requests": 15420,
"failed": 31,
"requestsPerMinute": 10.71,
"avgResponseTimeMs": 450.2,
"totalTokens": 2500000
},
"meta": { "requestId": "req_abc123" }
}

avgResponseTimeMs and totalTokens are null when no request in the window recorded them. When older records have been archived, the figures cover coveredSince to now. The figures count the requests you may count: a platform administrator every request, an organization owner or admin on a multi-tenant platform their organization's members', and anyone else their own (see Monitoring).


GET /api/v1/ai-gateway/models/:id/logs​

Get model logs

Returns container logs for a self-hosted model deployment.

Scope: ai-gateway:read

Parameters:

NameInTypeRequiredDescription
idpathstringYesModel ID
linesqueryintegerNoNumber of log lines to return (default: 100)
sincequerystringNoReturn logs since this time (ISO 8601 or duration like 1h, 30m)
containerquerystringNoContainer name (for multi-container pods)

Response: 200 OK

For a self-hosted model, data is the list of log entries:

{
"data": [
{ "timestamp": "2025-02-07T10:00:00Z", "level": "info", "message": "Model loaded successfully" },
{ "timestamp": "2025-02-07T10:00:01Z", "level": "info", "message": "Ready to serve requests on port 8000" }
],
"meta": { "requestId": "req_abc123" }
}

A third-party model has no deployment logs; data is { "logs": [], "message": "Third-party models do not generate deployment logs" }.


Sharing​

Who can reach a AI model is the platform's one sharing shape: its owner, its members (role editor can use and change it, user can only use it) and its visibility (public: every user can find and use it; in a multi-tenant deployment, every user of its organization). Changing it stays with its owner and editors.

MethodPathDoesScope
GET/api/v1/ai-gateway/models/:id/permissionsOwner, members (userId, role) and visibilityai-gateway:read
POST/api/v1/ai-gateway/models/:id/permissions/membersShare with a user: { "userId", "role": "editor" | "user" }ai-gateway:write
DELETE/api/v1/ai-gateway/models/:aiModelId/permissions/members/:userIdStop sharing with a userai-gateway:write
PATCH/api/v1/ai-gateway/models/:id/permissions{ "visibility": "public" | "private" }ai-gateway:write

Each, except a member's removal (204), answers the permissions as they are now:

{
"data": {
"resourceId": "<AI model id>",
"owner": "<user id>",
"members": [{ "userId": "<user id>", "role": "user" }],
"visibility": "private"
},
"meta": { "requestId": "req_abc123" }
}