Skip to main content

Model Registry

Register, manage, version, and deploy machine learning models. The model registry provides a centralized catalog for tracking models across your organization, regardless of their source or framework.

All endpoints require authentication via X-API-Key header and the appropriate scope.


RegisteredModel Object​

{
"_id": "model_abc123",
"name": "Churn Prediction XGBoost",
"description": "XGBoost model for customer churn prediction trained on Q4 2024 data",
"framework": "xgboost",
"algorithm": "XGBClassifier",
"tags": ["production", "churn", "xgboost"],
"source": "automl",
"workspaceId": "ws_001",
"organizationId": "org_xyz",
"artifact": { "s3Key": "model-artifacts/org_xyz/upload_abc123.zip", "sizeMb": 12.5 },
"features": { "featureNames": ["tenure", "monthly_spend", "support_tickets", "plan_type"], "targetName": "churned" },
"training": { "problemType": "classification" },
"metrics": { "classification": { "accuracy": 0.94 } },
"deployment": {
"status": "running",
"resources": { "cpu": "2", "memory": "4Gi", "replicas": 1 },
"deployedAt": "2025-02-01T10:00:00Z"
},
"access": { "owner": "user_456", "isPublic": false, "sharedWith": [] },
"versions": [
{
"version": 1,
"description": "Initial version",
"artifact": { "s3Key": "model-artifacts/org_xyz/upload_abc123.zip", "sizeMb": 12.5 },
"metrics": { "classification": { "accuracy": 0.94 } },
"deployment": { "status": "registered" },
"createdAt": "2025-01-15T10:30:00Z",
"createdBy": "user_456"
}
],
"activeVersion": 1,
"monitoring": { "recordInputs": true, "labelWindowDays": 14 },
"latestDrift": { "resultId": "6ab95770...", "overallStatus": "ok", "driftScore": 0.04, "sampleSize": 1250, "calculatedAt": "2026-09-27T00:05:12Z" },
"latestAnalysis": { "jobId": "6ab95740...", "status": "completed", "triggeredBy": "scheduler", "startedAt": "2026-09-27T00:00:03Z", "updatedAt": "2026-09-27T00:05:12Z", "finishedAt": "2026-09-27T00:05:12Z" },
"createdAt": "2025-01-15T10:30:00Z",
"updatedAt": "2025-02-01T10:00:00Z"
}
FieldDescription
sourceautoml, upload, experiment or external
deploymentThe serving state: status (registered, building, deploying, starting, running, idle, deployed, stopped, failed), resources, and deployedAt once deployed
accessowner (the owner's user ID), isPublic and sharedWith
versionsEvery version: its artifact, metrics, deployment and who created it when
activeVersionThe version the model serves, and the one drift analyzes
monitoringrecordInputs (keep each prediction's inputs and whole output; default on) and labelWindowDays (how long an actual keyed by entity id may take to arrive; default 14)
latestDriftThe latest drift result's summary (driftScore is the mean PSI)
latestAnalysisThe latest drift analysis run: status (pending, deploying, running, completed, failed), triggeredBy (manual, scheduler, baseline), statusMessage, and error saying why it produced no result

GET /api/v1/mlops/models​

List all registered models accessible to the authenticated user.

Scope: model-registry:read

Query Parameters

ParameterTypeRequiredDescription
qstringNoSearch by model name or description
frameworkstringNoFilter by framework: pytorch, tensorflow, sklearn, xgboost, lightgbm, onnx, custom, strongly-automl
sourcestringNoFilter by source: automl, upload, experiment, external
statusstringNoFilter by deployment status (deployment.status): registered, building, deploying, starting, running, idle, deployed, stopped, failed
tagstringNoFilter by tag
workspaceIdstringNoFilter by workspace ID
limitintegerNoNumber of results to return (default: 50, max: 200)
cursorstringNometa.nextCursor of the previous page; omit for the first page
sortstringNoSort field, prefixed with - for descending (default: -createdAt)

Response 200 OK

{
"data": [
{
"_id": "model_abc123",
"name": "Churn Prediction XGBoost",
"description": "XGBoost model for customer churn prediction trained on Q4 2024 data",
"framework": "xgboost",
"tags": ["production", "churn", "xgboost"],
"source": "automl",
"workspaceId": "ws_001",
"organizationId": "org_xyz",
"training": { "problemType": "classification" },
"deployment": { "status": "running" },
"access": { "owner": "user_456", "isPublic": false, "sharedWith": [] },
"activeVersion": 1,
"createdAt": "2025-01-15T10:30:00Z",
"updatedAt": "2025-02-01T10:00:00Z"
}
],
"meta": {
"total": 34,
"limit": 50,
"nextCursor": null,
"requestId": "req_abc123"
}
}

List items leave out versions, artifact and the build output; get a model by ID for those.


POST /api/v1/mlops/models/uploads​

A presigned link to upload a model bundle of up to 5 GB straight to storage. PUT the file to uploadUrl with Content-Length equal to fileSize, then register the model with the returned s3Key.

Scope: model-registry:write

FieldTypeRequiredDescription
fileNamestringYesThe bundle's file name, e.g. my-model.zip
fileSizenumberYesIts exact size in bytes

Response 201 Created: data is { "uploadUrl": "...", "s3Key": "...", "bucket": "...", "expiresIn": seconds }.


POST /api/v1/mlops/models/uploads/validate​

Upload a model artifact bundle to S3 and parse its manifest. Routes through the same upload path the UI uses, so manifest extraction, validation, and S3 upload behave identically. The returned s3Key + manifest are what POST /mlops/models (register) expects.

Scope: model-registry:write

Request (multipart/form-data)

FieldTypeRequiredDescription
filefileYesBundle binary -- either a .zip containing strongly.manifest.yaml at root (or one level nested), or a single .pkl / .joblib / .pt / .onnx artifact. Streamed up to 500 MB.

Response 201 Created

{
"data": {
"s3Key": "model-artifacts/org_xyz/upload_abc123.zip",
"bucket": "strongly-models",
"sizeMb": 12.5,
"manifest": {
"name": "Churn Prediction XGBoost",
"framework": "xgboost",
"version": "1.0.0",
"entrypoint": "predict.py"
},
"manifestRaw": "name: Churn Prediction XGBoost\nframework: xgboost\n..."
},
"meta": { "requestId": "req_abc123" }
}

The manifest / manifestRaw fields are only populated when the bundle is a .zip containing a valid strongly.manifest.yaml.


POST /api/v1/mlops/models​

Register a new model in the registry.

Scope: model-registry:write

Request Body

{
"name": "Churn Prediction XGBoost",
"framework": "xgboost",
"description": "XGBoost model for customer churn prediction trained on Q4 2024 data",
"source": "automl",
"tags": ["churn", "xgboost"],
"workspaceId": "ws_001",
"problemType": "classification",
"features": { "featureNames": ["tenure", "monthly_spend", "support_tickets", "plan_type"] }
}
FieldTypeRequiredDescription
namestringYesHuman-readable model name
frameworkstringYesModel framework: pytorch, tensorflow, xgboost, sklearn, onnx, custom
descriptionstringNoModel description
sourcestringNoWhere the model originated: automl, upload, experiment, external (default: external)
tagsstring[]NoArray of tag strings for categorization
workspaceIdstringNoWorkspace to register the model in
problemTypestringNoThe model's task: classification, regression, multiclass, multilabel, timeseries or other (any other value is refused). Stored at training.problemType; drift monitoring needs classification or regression
featuresobjectConditional{ "featureNames": [...] } -- the training feature columns in input order. Required for tabular frameworks (sklearn, xgboost, lightgbm, strongly-automl) so predictions are labelled for drift analysis
artifactobjectConditional{ s3Key, s3Bucket?, sizeMb? } from /model-registry/upload. Required for a servable model
manifestobjectConditionalThe parsed strongly.manifest.yaml a zip bundle's upload returned
frameworkVersionstringConditionalThe library version a single model file was saved with. Required for a single .joblib / .pkl / .pt file of sklearn, xgboost, lightgbm or pytorch: it is served with exactly that version
pythonVersionstringConditionalThe Python a single .joblib / .pkl file was saved with. Required for a pickle
schemaobjectConditional{ "input": { "fields": [{ "name", "type" }] } }: the model's inputs in the order it takes them (type float, int, string or bool). Required for a single model file: inputs are fed in this order
requirementsstring[]Nopip requirements a single model file's pickle imports beyond its framework (e.g. category-encoders==2.6.3), installed in its serving image
environmentVariablesobjectNoThe model's own environment variables { NAME: value }, set on its serving pod. Names the platform sets (MODEL_*, STRONGLY_*, AWS_*, FRAMEWORK, S3_BUCKET, INPUT_FIELDS, IMAGE_COLUMNS, PREDICT_PATH, HEALTH_PATH, PORT) are refused (400)

A single model file (.joblib, .pkl, .pickle, .pt, .pth, .onnx) is registered only with what serving it needs; a missing piece is refused (400) naming each. Its format decides its framework: a .joblib / .pkl is sklearn, xgboost, lightgbm or custom; a .pt is pytorch (TorchScript); a .onnx is onnx.

Response 201 Created

The new registered model.


GET /api/v1/mlops/models/:id​

Get a single registered model by ID, including all versions.

Scope: model-registry:read

Path Parameters

ParameterTypeRequiredDescription
idstringYesModel ID

Response 200 OK

Returns the full RegisteredModel object.


PATCH /api/v1/mlops/models/:id​

Update a registered model's metadata.

Scope: model-registry:write

Path Parameters

ParameterTypeRequiredDescription
idstringYesModel ID

Request Body

{
"name": "Churn Prediction XGBoost v2",
"description": "Updated model description with latest findings",
"tags": ["production", "churn", "xgboost", "v2"],
"problemType": "classification"
}

Only the fields sent change.

FieldTypeRequiredDescription
namestringNoModel name
descriptionstringNoModel description
tagsstring[]NoArray of tag strings
problemTypestringNoThe model's task: classification, regression, multiclass, multilabel, timeseries or other. Sets training.problemType and keeps the rest of the training record (and sets it on the active version). Drift monitoring needs classification or regression: set it on a model registered without one
trainingobjectNoThe whole training record (replaces it)

Response 200 OK

The updated registered model.


DELETE /api/v1/mlops/models/:id​

Delete a registered model. Deployed models must be undeployed before deletion. This action is irreversible.

Scope: model-registry:write

Path Parameters

ParameterTypeRequiredDescription
idstringYesModel ID

Response 204 No Content


POST /api/v1/mlops/models/:id/deploy​

Deploy a registered model for inference. Provisions the required infrastructure and creates a prediction endpoint.

Scope: model-registry:write

Path Parameters

ParameterTypeRequiredDescription
idstringYesModel ID

Deploys the model's active version. To serve another version, deploy it from Versions.

Request Body

{
"resources": { "cpu": "2", "memory": "4Gi", "replicas": 1 },
"autoShutdownMinutes": 15
}
FieldTypeRequiredDescription
resourcesobjectNoThe one size the model runs at: cpu, memory, gpu, gpuType, replicas (its container requests exactly what it is limited to)
autoShutdownMinutesnumberNoIdle minutes before the model scales to zero; > 0 deploys it on demand (it wakes on its first prediction), 0 or omitted keeps it always on
schedulingModestringNoalways_on, on_demand or scheduled
scheduleRulesarrayWith scheduledThe start and stop times: [{ "id", "action": "start" | "stop", "dayOfWeek": [0-6] (0 = Sunday), "time": "HH:MM", "timezone": IANA name }], at least one start and one stop
autoscalingobjectNoThe always-on HPA: enabled, minReplicas, maxReplicas, cpuThreshold, memoryThreshold. Only with always_on
useSpotbooleanNoRun on spot capacity
spotFallbackbooleanNoWith useSpot, fall back to on-demand capacity when spot is unavailable (default true); false requires spot
scalingModestringNofixed (default: wakes to resources.replicas) or auto (scales between minReplicas and maxReplicas with live demand)
minReplicasnumberNoThe awake minimum for scalingMode auto
maxReplicasnumberNoThe awake maximum for scalingMode auto
targetConcurrencynumberNoIn-flight predictions per replica that auto sizes for
instanceTypestringNoA node instance type to run on

Response 200 OK​

The model, as GET /mlops/models/:id shows it: deployment.status reports the deploy.


PATCH /api/v1/mlops/models/:id/serving​

Changes how the model is served (its Environment tab) without deploying it. Scheduling and demand scaling apply to a deployed model at once; resources and spot from its next deploy or start.

Scope: model-registry:write

Request Body: any of the deploy fields above, and environmentId (an environment to size the model by, or custom for resources). A strategy's settings are checked together: scheduled needs scheduleRules, on_demand an autoShutdownMinutes of at least 1, scalingMode auto needs on_demand with minReplicas, maxReplicas and targetConcurrency, and autoscaling is the always-on strategy's.

{
"schedulingMode": "scheduled",
"scheduleRules": [
{ "id": "weekday-start", "action": "start", "dayOfWeek": [1, 2, 3, 4, 5], "time": "09:00", "timezone": "America/New_York" },
{ "id": "weekday-stop", "action": "stop", "dayOfWeek": [1, 2, 3, 4, 5], "time": "17:00", "timezone": "America/New_York" }
],
"useSpot": true
}

Response 200 OK: the model.

POST /api/v1/mlops/models/:id/stop​

Stop a deployed model: its serving pod is removed and it moves to stopped (it reads stopping until then). A stopped model keeps its versions, monitoring and history; start it again, or deploy another version.

Scope: model-registry:write

Response 200 OK: the model, as GET /mlops/models/:id shows it, deployment.status stopping.


POST /api/v1/mlops/models/:id/start​

Start a stopped model at the size it was deployed with. It is starting until its pod is ready, then running; a start that cannot complete (its image cannot be pulled, its container cannot start) leaves it failed with the reason in its deployment logs, scaled back to zero. Poll the model's deployment.status.

Scope: model-registry:write

Response 200 OK: the model, as GET /mlops/models/:id shows it, deployment.status starting.


GET /api/v1/mlops/models/:id/status​

The model's lifecycle: status (its deployment's: registered when never deployed, building, deploying, starting, running, idle, waking, stopped or failed), canServe (whether predictions can be made now), activeVersion (the version that serves), versionCount and updatedAt.

Scope: model-registry:read


GET /api/v1/mlops/models/:id/logs​

The model's image build logs, or with logType=runtime its serving container's logs. A model that is not running has no container: asked for runtime logs, it answers its build logs with type build, requested runtime and a warning saying why.

Scope: model-registry:read

QueryTypeRequiredDescription
logTypestringNobuild (default) or runtime

Response 200 OK: data is { "logs": [...], "type": "build" | "runtime", "total": n }


POST /api/v1/mlops/models/:id/predict​

Run inference on a deployed model (deployment status running). The request is proxied to the model's pod through the AI Gateway, since model pods are internal to the cluster. Every prediction is recorded with the model version that made it.

Scope: model-registry:write

Path Parameters

ParameterTypeRequiredDescription
idstringYesModel ID

Request Body

{
"inputData": [[3.1, 4.2], [8.0, 9.1]],
"entityId": "order-1042",
"source": "production"
}

inputData is the feature payload forwarded verbatim to the model's serving endpoint. For the built-in tabular server pass the feature rows directly (a list of rows as above), a single { "feature": value, ... } row object, or { "instances": [ ...rows ] }. Do not wrap the rows in { "features": ... } -- the built-in server does not recognize that shape.

FieldTypeRequiredDescription
inputDataarray | objectYesFeature rows forwarded to the model's serving endpoint (rows array, single row object, or {"instances": rows})
entityIdstringNoYour id for what is predicted; actuals can join on it
sourcestringNoproduction (default) or test; test predictions are kept out of drift and A/B results
pathstringNoThe route on the model to call when it serves more than one (for example parse_text); omit it for the model's default

Response 200 OK

{
"data": {
"predictions": [1, 0],
"probabilities": [[0.18, 0.82], [0.74, 0.26]],
"latencyMs": 12,
"predictionId": "pZ8kQ2",
"entityId": "order-1042"
},
"meta": { "requestId": "req_abc123" }
}

predictions holds one entry per input row. probabilities is present for classifiers that expose predict_proba (one probability vector per row); it is omitted for regressors. A single-row request also returns single-element arrays.

Keep predictionId (or your entityId) to record the prediction's actual later with POST /mlops/models/:id/actuals. Returns 409 if the model is not deployed.

Versions​

A registered model accumulates immutable versions. Version 1 is created at registration; add a new version from a freshly uploaded artifact, then choose which one is served.

GET /api/v1/mlops/models/:id/versions​

List the model's version history and which one is active.

{
"data": {
"versions": [
{ "version": 1, "description": "Initial version", "artifact": { "s3Key": "..." }, "createdAt": "..." },
{ "version": 2, "description": "retrained on more data", "artifact": { "s3Key": "..." }, "createdAt": "..." }
],
"activeVersion": 2
},
"meta": { "requestId": "req_abc123" }
}

POST /api/v1/mlops/models/:id/versions​

Add a new version. Upload the artifact first via POST /mlops/models/uploads/validate, then pass its returned reference.

Body:

{
"artifact": {
"s3Key": "uploads/model-v2.joblib", "s3Bucket": "...", "sizeMb": 4.2,
"frameworkVersion": "1.6.1", "pythonVersion": "3.11.9"
},
"description": "retrained on more data"
}

Returns 201 Created with the new version, as the model's versions list it. artifact.s3Key is required. A single model file's version carries, on its artifact, the frameworkVersion and pythonVersion it was saved with and any requirements, as registering one does; schema gives its inputs when they differ from the model's.

POST /api/v1/mlops/models/:modelId/versions/:id/activate​

Point the model's served artifact/metrics at :version without deploying. data is the model, as GET /mlops/models/:id shows it, its activeVersion the version.

POST /api/v1/mlops/models/:modelId/versions/:id/deploy​

Set :version active and (re)deploy the serving pod with that version's artifact. Rolling back is just deploying an earlier version again.

Model Evaluation​

Monitoring settings, actuals (the true outcomes of predictions) and drift baselines. Drift analyses, results and settings are in the Drift Detection API. Files of any size (actuals, reference rows) are uploaded straight to storage: request an upload link, PUT the file to it with the returned contentType, then start the job that reads it.

Response 200 OK: the model, as GET /mlops/models/:id shows it: deployment.status reports the deploy.

PATCH /api/v1/mlops/models/:id/monitoring​

Change the model's monitoring settings. They apply from its next prediction.

Scope: model-registry:write

FieldTypeRequiredDescription
recordInputsbooleanNoKeep each prediction's inputs and whole output. Turn it off for sensitive data: predictions are still recorded, but input drift cannot be measured.
labelWindowDaysnumberNoHow long an actual keyed by entity id may take to arrive and still label that entity's predictions: 1 to 90 days

Send at least one. Response 200 OK: the updated registered model.

POST /api/v1/mlops/models/:id/actuals​

Record actuals for the model's predictions.

Scope: model-registry:write

{
"actuals": [
{ "predictionId": "pZ8kQ2", "actual": "churn" },
{ "entityId": "order-1042", "actual": 1, "actualAt": "2026-09-20T12:00:00Z" }
]
}

Each actual is keyed by predictionId (the id a prediction returned) or entityId (your id sent with the prediction). An entity id actual labels that entity's predictions made within the model's labelWindowDays before actualAt (else its upload time); either may arrive first. A second actual for the same key replaces the first.

Response 200 OK

{
"data": { "recorded": 1, "replaced": 1, "rejected": [], "batchId": "b_91c2" },
"meta": { "requestId": "req_abc123" }
}

rejected lists each row that could not be recorded: { "row": 0, "reason": "..." }.

POST /api/v1/mlops/models/:id/actuals/uploads​

A link to upload a CSV of actuals of any size: a predictionId or entityId column, an actual column, and optionally actualAt.

Scope: model-registry:write | Body: { "filename": "actuals.csv" }

Response 200 OK: data is { "uploadId": "...", "uploadUrl": "https://...", "contentType": "text/csv" }

POST /api/v1/mlops/models/:id/actuals/imports​

Import a CSV of actuals: an upload, or a file on a shared volume you can read. A job reads it.

Scope: model-registry:write

{ "source": { "type": "upload", "uploadId": "...", "filename": "actuals.csv" } }

or { "source": { "type": "volume", "volumeId": "...", "path": "outcomes/actuals.csv" } }.

Response 201 Created: the import, with status pending.

GET /api/v1/mlops/models/:id/actuals/imports​

The model's actuals imports, newest first: status (pending, running, completed, failed), rowsRead, recorded, replaced, rejected, and errorMessage when one failed.

Scope: model-registry:read | Sort: -createdAt (default), status, rowsRead, recorded, rejected

GET /api/v1/mlops/models/:modelId/actuals/imports/:id/rejections​

The rows an import could not record, each with its row (line in the file) and reason.

Scope: model-registry:read

POST /api/v1/mlops/models/:id/baselines/uploads​

A link to upload a CSV of a version's reference rows (the data it was trained or validated on) of any size.

Scope: model-registry:write | Body: { "filename": "reference.csv" }

POST /api/v1/mlops/models/:id/baselines​

Build a version's drift baseline from its reference rows: a column per input feature the model declares, and optionally actual, prediction and confidence (to record the version's accuracy, and estimate accuracy before actuals arrive).

Scope: model-registry:write

{
"source": { "type": "upload", "uploadId": "...", "filename": "reference.csv" },
"version": 2,
"analyzeDaily": true
}
FieldTypeRequiredDescription
sourceobjectYesAn upload (type: "upload", uploadId, filename) or a volume file (type: "volume", volumeId, path)
versionnumberNoThe version the rows describe (default: the active version). A model registered without versions has one baseline for the model: leave version out
analyzeDailybooleanNoTurn on the model's daily drift analysis once the baseline is kept (default false); the baseline's first analysis counts as that day's run

Response 201 Created: the job, with status pending. Once built, the baseline becomes the version's active one (the one it replaces is kept, inactive) and the version's first drift analysis starts, when the version has production predictions from the last 7 days (else the job's driftWaiting says so).

GET /api/v1/mlops/models/:id/baselines/jobs​

The model's baseline jobs, newest first: status (pending, running, completed, failed), modelVersion, sampleSize, baselineId, analyzeDaily, driftJobId (the first analysis), driftWaiting, driftError, driftScheduleError, and errorMessage when one failed.

Scope: model-registry:read | Sort: -createdAt (default), status, modelVersion

POST /api/v1/mlops/models/:modelId/baselines/:id/activate​

Make an earlier baseline its version's active one again; drift compares the version with it from the next analysis on.

Scope: model-registry:write

Response 200 OK: data is { "baselineId": "...", "modelVersion": 2, "isActive": true }

Usage analytics​

GET /api/v1/mlops/models/:id/analytics​

The model's usage over 7, 30 or 90 days, as its Analytics tab shows it: summary (predictions, errors, errorRate, callers, latencyAvgMs, latencyP50Ms, latencyP95Ms, avgConfidence, lastPredictionAt), daily (predictions, errors and latency per day), sources, versions and A/B variants (predictions and errors of each, with the test's abTestName and abTestDeleted when it has been deleted since), and weekdayHour (predictions by weekday, Monday first, and UTC hour).

Scope: ml-workbench:read

QueryTypeRequiredDescription
rangestringNo7d, 30d (default) or 90d

GET /api/v1/mlops/models/:id/analytics/callers​

Who called the model over the range, one row per caller and what the call came through (via: app, workflow, workspace or direct, with viaName): name, email, predictions, errors, lastSeen. Paged.

Scope: ml-workbench:read | Sort: -predictions (default), errors, lastSeen, name

QueryTypeRequiredDescription
rangestringNo7d, 30d (default) or 90d
qstringNoBy name, email, app or workflow
limit, cursornumberNoPaging