Skip to main content

Spot Instance Support

Use AWS Spot Instances to reduce compute costs by 60-90% for suitable workloads.

Overview​

Spot instances are spare EC2 capacity offered at significantly reduced prices. The trade-off is that AWS can reclaim them with 2 minutes notice when capacity is needed elsewhere.

Strongly AI supports spot instances through Karpenter NodePools. Spot is chosen on each workload, never on an environment: an environment is a size, and the workload you run at that size decides whether it runs on spot. If a spot instance is reclaimed, Kubernetes restarts the workload on a new node; persistent data (volumes) is not affected.

The spot setting​

Every workload that can run on spot has the same two settings:

  • Use Spot Instances (save up to 70%): runs the workload on spot capacity. Turning it on asks you to confirm the trade-offs, then shows a warning that the workload can be interrupted.
  • Fall back to on-demand when spot capacity is unavailable (recommended): shown once spot is on, and checked by default. Checked, spot is a scheduling preference: the workload runs on a spot node when one has room or a new node is added for it, and on on-demand capacity otherwise (including an on-demand node that already has room), so it is never held up waiting for spot. Unchecked, the workload runs on spot only and waits until spot capacity exists.

Detail pages show the choice as Capacity: On-demand, Spot (falls back to on-demand) or Spot only, and so do the create forms' summaries. Where a running workload's size can be changed (an app's Size card, a self-hosted model's Environment tab, a registry model's Environment tab, a job's Edit), its capacity is changed in the same place; where the size is fixed at create (workspaces, avatars, AutoML and fine-tuning jobs), so is the capacity.

Where it appears​

WorkloadWhere you set itREST fields
AppsDeploy App, the marketplace deploy's App Resources step, and the app's Size card (Capacity, Edit)useSpot, spotFallback on POST /apps and PATCH /apps/:id (from the next deploy, start or restart); config.resources.useSpot, config.resources.spotFallback on POST /marketplace/deploy
WorkspacesCreate WorkspaceuseSpot, spotFallback on POST /workspaces
Project jobsCreate Job, and the job's EdituseSpot, spotFallback on POST /jobs and PATCH /jobs/:id (from the next run)
Avatars (GPU tiers)Create Avatarresources.useSpot, resources.spotFallback
Self-hosted modelsDeploy form, Resources step; the model's Environment tab (Edit Configuration), which moves the running model on or off spotuseSpot, spotFallback on the deploy and on PATCH /ai-gateway/models/:id/resources (the model restarts on the new capacity)
Fine-tuned modelsDeploy Fine-Tuned Model dialoguseSpot, spotFallback on POST /ai-gateway/fine-tuning-jobs/:id/deploy
Model Registry modelsThe model's Environment tabuseSpot, spotFallback on deploy and PATCH /mlops/models/:id/serving
AutoML jobsCreate AutoML Job, Hardware stephardware.useSpot, hardware.spotFallback
Fine-tuning jobsCreate Fine-Tuning Job, Resourcesenvironment.useSpot, environment.spotFallback

spotFallback defaults to true everywhere; it only matters when useSpot is true.

# An app on spot, falling back to on-demand (the default)
POST /api/v1/apps
{ "name": "...", "cpu": "1", "memory": "2GB", "useSpot": true }

# A workspace on spot only
POST /api/v1/workspaces
{ "name": "...", "useSpot": true, "spotFallback": false, ... }

Workload Compatibility​

WorkloadSpot?Why
AppsYes (opt-in)Restart on a new node; volumes are kept. In-memory state is lost
WorkspacesYes (opt-in)Volumes are kept; open notebooks, terminals and processes are lost
Project jobsYes (opt-in)A reclaimed run is interrupted; suited to runs that can be repeated
Fine-tuning jobsYes (opt-in)Training checkpoints let an interrupted job resume
AutoML jobsYes (opt-in)Batch jobs that can be re-run
AvatarsYes (opt-in)A reclaim drops a live call; suited to development avatars
Self-hosted and registry modelsYes (opt-in)Stateless inference; callers retry
WorkflowsNoWorkflow workers are pinned to on-demand nodes
Add-onsNoDatabases run on the dedicated user-addons pool (no spot, never moved)

Interruptions​

When a spot node is reclaimed, AWS gives two minutes' notice and Kubernetes reschedules the pods on a replacement node: on spot when spot capacity exists, or on on-demand when the workload falls back. In-flight requests to a model or app fail and are retried by their callers; a job's run is interrupted.

Admin Controls​

Admins control spot availability through the Compute admin page (Admin > Compute, /admin/compute):

  • Enable or disable the "Spot (60-90% cheaper, can be interrupted)" capacity type per workload pool. The default for every pool is on-demand only.
  • Set resource limits per pool (pools are uncapped by default)

A workload only lands on spot when both the admin has enabled the spot capacity type for its pool and the deployment itself opted into spot.