Model Registry
Register, manage, version, and deploy machine learning models. The model registry provides a centralized catalog for tracking models across your organization, regardless of their source or framework.
All endpoints require authentication via X-API-Key header and the appropriate scope.
RegisteredModel Object
{
"_id": "model_abc123",
"name": "Churn Prediction XGBoost",
"description": "XGBoost model for customer churn prediction trained on Q4 2024 data",
"framework": "xgboost",
"algorithm": "XGBClassifier",
"tags": ["production", "churn", "xgboost"],
"source": "automl",
"workspaceId": "ws_001",
"organizationId": "org_xyz",
"artifact": { "s3Key": "model-artifacts/org_xyz/upload_abc123.zip", "sizeMb": 12.5 },
"features": { "featureNames": ["tenure", "monthly_spend", "support_tickets", "plan_type"], "targetName": "churned" },
"training": { "problemType": "classification" },
"metrics": { "classification": { "accuracy": 0.94 } },
"deployment": {
"status": "running",
"resources": { "cpu": "2", "memory": "4Gi", "replicas": 1 },
"deployedAt": "2025-02-01T10:00:00Z"
},
"access": { "owner": "user_456", "isPublic": false, "sharedWith": [] },
"versions": [
{
"version": 1,
"description": "Initial version",
"artifact": { "s3Key": "model-artifacts/org_xyz/upload_abc123.zip", "sizeMb": 12.5 },
"metrics": { "classification": { "accuracy": 0.94 } },
"deployment": { "status": "registered" },
"createdAt": "2025-01-15T10:30:00Z",
"createdBy": "user_456"
}
],
"activeVersion": 1,
"monitoring": { "recordInputs": true, "labelWindowDays": 14 },
"latestDrift": { "resultId": "6ab95770...", "overallStatus": "ok", "driftScore": 0.04, "sampleSize": 1250, "calculatedAt": "2026-09-27T00:05:12Z" },
"latestAnalysis": { "jobId": "6ab95740...", "status": "completed", "triggeredBy": "scheduler", "startedAt": "2026-09-27T00:00:03Z", "updatedAt": "2026-09-27T00:05:12Z", "finishedAt": "2026-09-27T00:05:12Z" },
"createdAt": "2025-01-15T10:30:00Z",
"updatedAt": "2025-02-01T10:00:00Z"
}
| Field | Description |
|---|---|
source | automl, upload, experiment or external |
deployment | The serving state: status (registered, building, deploying, starting, running, idle, deployed, stopped, failed), resources, and deployedAt once deployed |
access | owner (the owner's user ID), isPublic and sharedWith |
versions | Every version: its artifact, metrics, deployment and who created it when |
activeVersion | The version the model serves, and the one drift analyzes |
monitoring | recordInputs (keep each prediction's inputs and whole output; default on) and labelWindowDays (how long an actual keyed by entity id may take to arrive; default 14) |
latestDrift | The latest drift result's summary (driftScore is the mean PSI) |
latestAnalysis | The latest drift analysis run: status (pending, deploying, running, completed, failed), triggeredBy (manual, scheduler, baseline), statusMessage, and error saying why it produced no result |
GET /api/v1/mlops/models
List all registered models accessible to the authenticated user.
Scope: model-registry:read
Query Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
q | string | No | Search by model name or description |
framework | string | No | Filter by framework: pytorch, tensorflow, sklearn, xgboost, lightgbm, onnx, custom, strongly-automl |
source | string | No | Filter by source: automl, upload, experiment, external |
status | string | No | Filter by deployment status (deployment.status): registered, building, deploying, starting, running, idle, deployed, stopped, failed |
tag | string | No | Filter by tag |
workspaceId | string | No | Filter by workspace ID |
limit | integer | No | Number of results to return (default: 50, max: 200) |
cursor | string | No | meta.nextCursor of the previous page; omit for the first page |
sort | string | No | Sort field, prefixed with - for descending (default: -createdAt) |
Response 200 OK
{
"data": [
{
"_id": "model_abc123",
"name": "Churn Prediction XGBoost",
"description": "XGBoost model for customer churn prediction trained on Q4 2024 data",
"framework": "xgboost",
"tags": ["production", "churn", "xgboost"],
"source": "automl",
"workspaceId": "ws_001",
"organizationId": "org_xyz",
"training": { "problemType": "classification" },
"deployment": { "status": "running" },
"access": { "owner": "user_456", "isPublic": false, "sharedWith": [] },
"activeVersion": 1,
"createdAt": "2025-01-15T10:30:00Z",
"updatedAt": "2025-02-01T10:00:00Z"
}
],
"meta": {
"total": 34,
"limit": 50,
"nextCursor": null,
"requestId": "req_abc123"
}
}
List items leave out versions, artifact and the build output; get a model by ID for those.
POST /api/v1/mlops/models/uploads
A presigned link to upload a model bundle of up to 5 GB straight to storage. PUT the file to uploadUrl with Content-Length equal to fileSize, then register the model with the returned s3Key.
Scope: model-registry:write
| Field | Type | Required | Description |
|---|---|---|---|
fileName | string | Yes | The bundle's file name, e.g. my-model.zip |
fileSize | number | Yes | Its exact size in bytes |
Response 201 Created: data is { "uploadUrl": "...", "s3Key": "...", "bucket": "...", "expiresIn": seconds }.
POST /api/v1/mlops/models/uploads/validate
Upload a model artifact bundle to S3 and parse its manifest. Routes through the same upload path the UI uses, so manifest extraction, validation, and S3 upload behave identically. The returned s3Key + manifest are what POST /mlops/models (register) expects.
Scope: model-registry:write
Request (multipart/form-data)
| Field | Type | Required | Description |
|---|---|---|---|
file | file | Yes | Bundle binary -- either a .zip containing strongly.manifest.yaml at root (or one level nested), or a single .pkl / .joblib / .pt / .onnx artifact. Streamed up to 500 MB. |
Response 201 Created
{
"data": {
"s3Key": "model-artifacts/org_xyz/upload_abc123.zip",
"bucket": "strongly-models",
"sizeMb": 12.5,
"manifest": {
"name": "Churn Prediction XGBoost",
"framework": "xgboost",
"version": "1.0.0",
"entrypoint": "predict.py"
},
"manifestRaw": "name: Churn Prediction XGBoost\nframework: xgboost\n..."
},
"meta": { "requestId": "req_abc123" }
}
The manifest / manifestRaw fields are only populated when the bundle is a .zip containing a valid strongly.manifest.yaml.
POST /api/v1/mlops/models
Register a new model in the registry.
Scope: model-registry:write
Request Body
{
"name": "Churn Prediction XGBoost",
"framework": "xgboost",
"description": "XGBoost model for customer churn prediction trained on Q4 2024 data",
"source": "automl",
"tags": ["churn", "xgboost"],
"workspaceId": "ws_001",
"problemType": "classification",
"features": { "featureNames": ["tenure", "monthly_spend", "support_tickets", "plan_type"] }
}
| Field | Type | Required | Description |
|---|---|---|---|
name | string | Yes | Human-readable model name |
framework | string | Yes | Model framework: pytorch, tensorflow, xgboost, sklearn, onnx, custom |
description | string | No | Model description |
source | string | No | Where the model originated: automl, upload, experiment, external (default: external) |
tags | string[] | No | Array of tag strings for categorization |
workspaceId | string | No | Workspace to register the model in |
problemType | string | No | The model's task: classification, regression, multiclass, multilabel, timeseries or other (any other value is refused). Stored at training.problemType; drift monitoring needs classification or regression |
features | object | Conditional | { "featureNames": [...] } -- the training feature columns in input order. Required for tabular frameworks (sklearn, xgboost, lightgbm, strongly-automl) so predictions are labelled for drift analysis |
artifact | object | Conditional | { s3Key, s3Bucket?, sizeMb? } from /model-registry/upload. Required for a servable model |
manifest | object | Conditional | The parsed strongly.manifest.yaml a zip bundle's upload returned |
frameworkVersion | string | Conditional | The library version a single model file was saved with. Required for a single .joblib / .pkl / .pt file of sklearn, xgboost, lightgbm or pytorch: it is served with exactly that version |
pythonVersion | string | Conditional | The Python a single .joblib / .pkl file was saved with. Required for a pickle |
schema | object | Conditional | { "input": { "fields": [{ "name", "type" }] } }: the model's inputs in the order it takes them (type float, int, string or bool). Required for a single model file: inputs are fed in this order |
requirements | string[] | No | pip requirements a single model file's pickle imports beyond its framework (e.g. category-encoders==2.6.3), installed in its serving image |
environmentVariables | object | No | The model's own environment variables { NAME: value }, set on its serving pod. Names the platform sets (MODEL_*, STRONGLY_*, AWS_*, FRAMEWORK, S3_BUCKET, INPUT_FIELDS, IMAGE_COLUMNS, PREDICT_PATH, HEALTH_PATH, PORT) are refused (400) |
A single model file (.joblib, .pkl, .pickle, .pt, .pth, .onnx) is registered only with what serving it needs; a missing piece is refused (400) naming each. Its format decides its framework: a .joblib / .pkl is sklearn, xgboost, lightgbm or custom; a .pt is pytorch (TorchScript); a .onnx is onnx.
Response 201 Created
The new registered model.
GET /api/v1/mlops/models/:id
Get a single registered model by ID, including all versions.
Scope: model-registry:read
Path Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
id | string | Yes | Model ID |
Response 200 OK
Returns the full RegisteredModel object.
PATCH /api/v1/mlops/models/:id
Update a registered model's metadata.
Scope: model-registry:write
Path Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
id | string | Yes | Model ID |
Request Body
{
"name": "Churn Prediction XGBoost v2",
"description": "Updated model description with latest findings",
"tags": ["production", "churn", "xgboost", "v2"],
"problemType": "classification"
}
Only the fields sent change.
| Field | Type | Required | Description |
|---|---|---|---|
name | string | No | Model name |
description | string | No | Model description |
tags | string[] | No | Array of tag strings |
problemType | string | No | The model's task: classification, regression, multiclass, multilabel, timeseries or other. Sets training.problemType and keeps the rest of the training record (and sets it on the active version). Drift monitoring needs classification or regression: set it on a model registered without one |
training | object | No | The whole training record (replaces it) |
Response 200 OK
The updated registered model.
DELETE /api/v1/mlops/models/:id
Delete a registered model. Deployed models must be undeployed before deletion. This action is irreversible.
Scope: model-registry:write
Path Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
id | string | Yes | Model ID |
Response 204 No Content
POST /api/v1/mlops/models/:id/deploy
Deploy a registered model for inference. Provisions the required infrastructure and creates a prediction endpoint.
Scope: model-registry:write
Path Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
id | string | Yes | Model ID |
Deploys the model's active version. To serve another version, deploy it from Versions.
Request Body
{
"resources": { "cpu": "2", "memory": "4Gi", "replicas": 1 },
"autoShutdownMinutes": 15
}
| Field | Type | Required | Description |
|---|---|---|---|
resources | object | No | The one size the model runs at: cpu, memory, gpu, gpuType, replicas (its container requests exactly what it is limited to) |
autoShutdownMinutes | number | No | Idle minutes before the model scales to zero; > 0 deploys it on demand (it wakes on its first prediction), 0 or omitted keeps it always on |
schedulingMode | string | No | always_on, on_demand or scheduled |
scheduleRules | array | With scheduled | The start and stop times: [{ "id", "action": "start" | "stop", "dayOfWeek": [0-6] (0 = Sunday), "time": "HH:MM", "timezone": IANA name }], at least one start and one stop |
autoscaling | object | No | The always-on HPA: enabled, minReplicas, maxReplicas, cpuThreshold, memoryThreshold. Only with always_on |
useSpot | boolean | No | Run on spot capacity |
spotFallback | boolean | No | With useSpot, fall back to on-demand capacity when spot is unavailable (default true); false requires spot |
scalingMode | string | No | fixed (default: wakes to resources.replicas) or auto (scales between minReplicas and maxReplicas with live demand) |
minReplicas | number | No | The awake minimum for scalingMode auto |
maxReplicas | number | No | The awake maximum for scalingMode auto |
targetConcurrency | number | No | In-flight predictions per replica that auto sizes for |
instanceType | string | No | A node instance type to run on |
Response 200 OK
The model, as GET /mlops/models/:id shows it: deployment.status reports the deploy.
PATCH /api/v1/mlops/models/:id/serving
Changes how the model is served (its Environment tab) without deploying it. Scheduling and demand scaling apply to a deployed model at once; resources and spot from its next deploy or start.
Scope: model-registry:write
Request Body: any of the deploy fields above, and environmentId (an environment to size the model by, or custom for resources). A strategy's settings are checked together: scheduled needs scheduleRules, on_demand an autoShutdownMinutes of at least 1, scalingMode auto needs on_demand with minReplicas, maxReplicas and targetConcurrency, and autoscaling is the always-on strategy's.
{
"schedulingMode": "scheduled",
"scheduleRules": [
{ "id": "weekday-start", "action": "start", "dayOfWeek": [1, 2, 3, 4, 5], "time": "09:00", "timezone": "America/New_York" },
{ "id": "weekday-stop", "action": "stop", "dayOfWeek": [1, 2, 3, 4, 5], "time": "17:00", "timezone": "America/New_York" }
],
"useSpot": true
}
Response 200 OK: the model.
POST /api/v1/mlops/models/:id/stop
Stop a deployed model: its serving pod is removed and it moves to stopped (it reads stopping until then). A stopped model keeps its versions, monitoring and history; start it again, or deploy another version.
Scope: model-registry:write
Response 200 OK: the model, as GET /mlops/models/:id shows it, deployment.status stopping.
POST /api/v1/mlops/models/:id/start
Start a stopped model at the size it was deployed with. It is starting until its pod is ready, then running; a start that cannot complete (its image cannot be pulled, its container cannot start) leaves it failed with the reason in its deployment logs, scaled back to zero. Poll the model's deployment.status.
Scope: model-registry:write
Response 200 OK: the model, as GET /mlops/models/:id shows it, deployment.status starting.
GET /api/v1/mlops/models/:id/status
The model's lifecycle: status (its deployment's: registered when never deployed, building, deploying, starting, running, idle, waking, stopped or failed), canServe (whether predictions can be made now), activeVersion (the version that serves), versionCount and updatedAt.
Scope: model-registry:read
GET /api/v1/mlops/models/:id/logs
The model's image build logs, or with logType=runtime its serving container's logs. A model that is not running has no container: asked for runtime logs, it answers its build logs with type build, requested runtime and a warning saying why.
Scope: model-registry:read
| Query | Type | Required | Description |
|---|---|---|---|
logType | string | No | build (default) or runtime |
Response 200 OK: data is { "logs": [...], "type": "build" | "runtime", "total": n }
POST /api/v1/mlops/models/:id/predict
Run inference on a deployed model (deployment status running). The request
is proxied to the model's pod through the AI Gateway, since model pods are
internal to the cluster. Every prediction is recorded with the model version
that made it.
Scope: model-registry:write
Path Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
id | string | Yes | Model ID |
Request Body
{
"inputData": [[3.1, 4.2], [8.0, 9.1]],
"entityId": "order-1042",
"source": "production"
}
inputData is the feature payload forwarded verbatim to the model's serving
endpoint. For the built-in tabular server pass the feature rows directly (a list
of rows as above), a single { "feature": value, ... } row object, or
{ "instances": [ ...rows ] }. Do not wrap the rows in { "features": ... }
-- the built-in server does not recognize that shape.
| Field | Type | Required | Description |
|---|---|---|---|
inputData | array | object | Yes | Feature rows forwarded to the model's serving endpoint (rows array, single row object, or {"instances": rows}) |
entityId | string | No | Your id for what is predicted; actuals can join on it |
source | string | No | production (default) or test; test predictions are kept out of drift and A/B results |
path | string | No | The route on the model to call when it serves more than one (for example parse_text); omit it for the model's default |
Response 200 OK
{
"data": {
"predictions": [1, 0],
"probabilities": [[0.18, 0.82], [0.74, 0.26]],
"latencyMs": 12,
"predictionId": "pZ8kQ2",
"entityId": "order-1042"
},
"meta": { "requestId": "req_abc123" }
}
predictions holds one entry per input row. probabilities is present for
classifiers that expose predict_proba (one probability vector per row); it is
omitted for regressors. A single-row request also returns single-element arrays.
Keep predictionId (or your entityId) to record the prediction's actual later
with POST /mlops/models/:id/actuals.
Returns 409 if the model is not deployed.
Versions
A registered model accumulates immutable versions. Version 1 is created at registration; add a new version from a freshly uploaded artifact, then choose which one is served.
GET /api/v1/mlops/models/:id/versions
List the model's version history and which one is active.
{
"data": {
"versions": [
{ "version": 1, "description": "Initial version", "artifact": { "s3Key": "..." }, "createdAt": "..." },
{ "version": 2, "description": "retrained on more data", "artifact": { "s3Key": "..." }, "createdAt": "..." }
],
"activeVersion": 2
},
"meta": { "requestId": "req_abc123" }
}
POST /api/v1/mlops/models/:id/versions
Add a new version. Upload the artifact first via POST /mlops/models/uploads/validate,
then pass its returned reference.
Body:
{
"artifact": {
"s3Key": "uploads/model-v2.joblib", "s3Bucket": "...", "sizeMb": 4.2,
"frameworkVersion": "1.6.1", "pythonVersion": "3.11.9"
},
"description": "retrained on more data"
}
Returns 201 Created with the new version, as the model's versions list it. artifact.s3Key is required. A single model file's version carries, on its artifact, the frameworkVersion and pythonVersion it was saved with and any requirements, as registering one does; schema gives its inputs when they differ from the model's.
POST /api/v1/mlops/models/:modelId/versions/:id/activate
Point the model's served artifact/metrics at :version without deploying.
data is the model, as GET /mlops/models/:id shows it, its activeVersion the version.
POST /api/v1/mlops/models/:modelId/versions/:id/deploy
Set :version active and (re)deploy the serving pod with that version's
artifact. Rolling back is just deploying an earlier version again.
Model Evaluation
Monitoring settings, actuals (the true outcomes of predictions) and drift baselines. Drift analyses, results and settings are in the Drift Detection API. Files of any size (actuals, reference rows) are uploaded straight to storage: request an upload link, PUT the file to it with the returned contentType, then start the job that reads it.
Response 200 OK: the model, as GET /mlops/models/:id shows it: deployment.status reports the deploy.
PATCH /api/v1/mlops/models/:id/monitoring
Change the model's monitoring settings. They apply from its next prediction.
Scope: model-registry:write
| Field | Type | Required | Description |
|---|---|---|---|
recordInputs | boolean | No | Keep each prediction's inputs and whole output. Turn it off for sensitive data: predictions are still recorded, but input drift cannot be measured. |
labelWindowDays | number | No | How long an actual keyed by entity id may take to arrive and still label that entity's predictions: 1 to 90 days |
Send at least one. Response 200 OK: the updated registered model.
POST /api/v1/mlops/models/:id/actuals
Record actuals for the model's predictions.
Scope: model-registry:write
{
"actuals": [
{ "predictionId": "pZ8kQ2", "actual": "churn" },
{ "entityId": "order-1042", "actual": 1, "actualAt": "2026-09-20T12:00:00Z" }
]
}
Each actual is keyed by predictionId (the id a prediction returned) or entityId (your id sent with the prediction). An entity id actual labels that entity's predictions made within the model's labelWindowDays before actualAt (else its upload time); either may arrive first. A second actual for the same key replaces the first.
Response 200 OK
{
"data": { "recorded": 1, "replaced": 1, "rejected": [], "batchId": "b_91c2" },
"meta": { "requestId": "req_abc123" }
}
rejected lists each row that could not be recorded: { "row": 0, "reason": "..." }.
POST /api/v1/mlops/models/:id/actuals/uploads
A link to upload a CSV of actuals of any size: a predictionId or entityId column, an actual column, and optionally actualAt.
Scope: model-registry:write | Body: { "filename": "actuals.csv" }
Response 200 OK: data is { "uploadId": "...", "uploadUrl": "https://...", "contentType": "text/csv" }
POST /api/v1/mlops/models/:id/actuals/imports
Import a CSV of actuals: an upload, or a file on a shared volume you can read. A job reads it.
Scope: model-registry:write
{ "source": { "type": "upload", "uploadId": "...", "filename": "actuals.csv" } }
or { "source": { "type": "volume", "volumeId": "...", "path": "outcomes/actuals.csv" } }.
Response 201 Created: the import, with status pending.
GET /api/v1/mlops/models/:id/actuals/imports
The model's actuals imports, newest first: status (pending, running, completed, failed), rowsRead, recorded, replaced, rejected, and errorMessage when one failed.
Scope: model-registry:read | Sort: -createdAt (default), status, rowsRead, recorded, rejected
GET /api/v1/mlops/models/:modelId/actuals/imports/:id/rejections
The rows an import could not record, each with its row (line in the file) and reason.
Scope: model-registry:read
POST /api/v1/mlops/models/:id/baselines/uploads
A link to upload a CSV of a version's reference rows (the data it was trained or validated on) of any size.
Scope: model-registry:write | Body: { "filename": "reference.csv" }
POST /api/v1/mlops/models/:id/baselines
Build a version's drift baseline from its reference rows: a column per input feature the model declares, and optionally actual, prediction and confidence (to record the version's accuracy, and estimate accuracy before actuals arrive).
Scope: model-registry:write
{
"source": { "type": "upload", "uploadId": "...", "filename": "reference.csv" },
"version": 2,
"analyzeDaily": true
}
| Field | Type | Required | Description |
|---|---|---|---|
source | object | Yes | An upload (type: "upload", uploadId, filename) or a volume file (type: "volume", volumeId, path) |
version | number | No | The version the rows describe (default: the active version). A model registered without versions has one baseline for the model: leave version out |
analyzeDaily | boolean | No | Turn on the model's daily drift analysis once the baseline is kept (default false); the baseline's first analysis counts as that day's run |
Response 201 Created: the job, with status pending. Once built, the baseline becomes the version's active one (the one it replaces is kept, inactive) and the version's first drift analysis starts, when the version has production predictions from the last 7 days (else the job's driftWaiting says so).
GET /api/v1/mlops/models/:id/baselines/jobs
The model's baseline jobs, newest first: status (pending, running, completed, failed), modelVersion, sampleSize, baselineId, analyzeDaily, driftJobId (the first analysis), driftWaiting, driftError, driftScheduleError, and errorMessage when one failed.
Scope: model-registry:read | Sort: -createdAt (default), status, modelVersion
POST /api/v1/mlops/models/:modelId/baselines/:id/activate
Make an earlier baseline its version's active one again; drift compares the version with it from the next analysis on.
Scope: model-registry:write
Response 200 OK: data is { "baselineId": "...", "modelVersion": 2, "isActive": true }
Usage analytics
GET /api/v1/mlops/models/:id/analytics
The model's usage over 7, 30 or 90 days, as its Analytics tab shows it: summary (predictions, errors, errorRate, callers, latencyAvgMs, latencyP50Ms, latencyP95Ms, avgConfidence, lastPredictionAt), daily (predictions, errors and latency per day), sources, versions and A/B variants (predictions and errors of each, with the test's abTestName and abTestDeleted when it has been deleted since), and weekdayHour (predictions by weekday, Monday first, and UTC hour).
Scope: ml-workbench:read
| Query | Type | Required | Description |
|---|---|---|---|
range | string | No | 7d, 30d (default) or 90d |
GET /api/v1/mlops/models/:id/analytics/callers
Who called the model over the range, one row per caller and what the call came through (via: app, workflow, workspace or direct, with viaName): name, email, predictions, errors, lastSeen. Paged.
Scope: ml-workbench:read | Sort: -predictions (default), errors, lastSeen, name
| Query | Type | Required | Description |
|---|---|---|---|
range | string | No | 7d, 30d (default) or 90d |
q | string | No | By name, email, app or workflow |
limit, cursor | number | No | Paging |