Deploying AI Models
Deploy open-source models from Hugging Face on your own infrastructure with full control over scaling, costs, and data privacy. Deployments are scoped to organization namespaces for multi-tenant isolation.
Environment
Self-hosted models use one of several inference engines depending on the model type:
- vLLM (113 models) -- High-performance inference with PagedAttention, continuous batching, and OpenAI-compatible API. Used for LLMs, multimodal, and embedding models.
- Transformers (24 models) -- Custom Python servers using HuggingFace Transformers directly. Used for models not yet supported by vLLM.
- TTS Engine (16 models) -- Custom text-to-speech servers exposing
/v1/audio/speech. - Whisper Engine (5 models) -- Whisper-compatible STT servers exposing
/v1/audio/transcriptions. - Video Engine (6 models) -- Video generation servers exposing
/v1/video/generations. - Custom (5 models) -- Generic custom servers with configurable endpoints.
See Self-Hosted Models for the full engine architecture, provider routing, and model catalogue.
Multi-Tenant Deployment
Every model deployment belongs to your organization:
- It runs in your organization's own namespace, isolated from other organizations
- It is billed to your organization and checked against its budgets
- Governance requirements that apply to models must be met before it deploys
Deployment Options
Choose when your model runs in the deploy form's Scheduling step (or with the deploy fields over REST):
Always On
Behavior: The model runs 24/7 with the number of Replicas you set.
- Use Case: Production applications with consistent traffic
- Cost: Fixed compute costs regardless of usage
- Latency: Instant response (no cold start)
- Variable load: turn on Autoscaling to scale between a minimum and maximum on CPU and memory
Always On deployment ensures instant response times and is ideal for production applications with consistent traffic patterns.
On Demand (Auto-Shutdown)
Behavior: The model scales to zero after it has been idle for Auto-shutdown after (minutes), and wakes on the next request.
- Use Case: Development, testing, low-traffic applications, models used only inside workflow runs
- Cost: Pay only while the model is awake
- Latency: The first request after idle waits while the model starts
- Configuration: an idle window of at least 1 minute is required (
auto_shutdown_minutesover REST); the form refuses 0
On Demand deployment can significantly reduce costs for intermittent workloads by stopping the model during idle periods.
Demand scaling (On Demand + Auto-scaling)
An on-demand model can also scale with its traffic while it is awake. Choose Auto-scaling and set:
| Setting | REST field | Meaning |
|---|---|---|
| Min replicas | min_replicas | Fewest replicas while awake (at least 1) |
| Max replicas | max_replicas | Most replicas (at least Min replicas) |
| Target concurrency | target_concurrency | Requests in flight one replica handles before another is added |
While awake, the model adds replicas as soon as load rises and removes them only after load stays low, so it does not flap. When idle for the auto-shutdown window it still scales to zero. All three settings are required, and the scaling mode is exactly Fixed (fixed, always wake to Replicas) or Auto (auto); any other value is refused. Auto-scaling is only for on-demand models; always-on models use Autoscaling.
Scheduled
Behavior: The model starts and stops at the times you set and runs Replicas pods in between.
- Use Case: Batch processing, business hours only, regional availability
- Cost: Only charged while the schedule has it running
- Latency: Instant during scheduled hours, unavailable outside them
- Configuration: pick the days, a start time and a stop time, and a timezone. A schedule needs at least one start and one stop. Over REST, send
schedule_rules(see POST /api/v1/ai-gateway/models/:id/deploy).
Resource Configuration
Choose the right hardware for your model size:
| Model Size | GPU | VRAM | Recommended For |
|---|---|---|---|
| Small (7B parameters) | 1x V100 / A10G | 16-24GB | Llama-2-7B, Mistral-7B |
| Medium (13B parameters) | 1x A10G | 24GB | Llama-2-13B, Vicuna-13B |
| Large (70B parameters) | 4x A100 (40GB) | 160GB total | Llama-2-70B with tensor parallelism |
| Quantized (4-bit) | 1x V100 | 8-12GB | 70B models quantized to 4-bit |
Available Instance Types
The deployment system offers these EC2 instance types:
GPU Instances:
| Instance | GPU | GPU Memory | vCPUs | RAM |
|---|---|---|---|---|
g5.xlarge | 1x A10G | 24GB | 4 | 16GB |
g5.2xlarge | 1x A10G | 24GB | 8 | 32GB |
g5.12xlarge | 4x A10G | 96GB | 48 | 192GB |
p3.2xlarge | 1x V100 | 16GB | 8 | 61GB |
p3.8xlarge | 4x V100 | 64GB | 32 | 244GB |
CPU-Only Instances:
| Instance | vCPUs | RAM |
|---|---|---|
c5.2xlarge | 8 | 16GB |
c5.4xlarge | 16 | 32GB |
What a deployment costs is the infrastructure it runs on, shown in FinOps.
Capacity Options
Spot Instances
Self-hosted models can opt into AWS Spot capacity -- the same option available for AutoML and fine-tuning jobs.
UI: in the deploy form (Resources step), toggle Use Spot Instances. When enabled, the Fall back to on-demand when spot capacity is unavailable checkbox (recommended, checked by default) controls the behaviour:
- Checked (default): spot is requested as a scheduling preference: the model runs on a spot node when one has room or a new node is added for it, and on on-demand capacity otherwise. Availability is never sacrificed for price.
- Unchecked (strict spot): the deployment pins to spot and waits until spot capacity is available.
REST: pass use_spot / spot_fallback in the deploy body -- see
POST /api/v1/ai-gateway/models/:id/deploy and
Spot Instance Support.
If a spot node is reclaimed, the platform receives the 2-minute EC2 warning, drains the node gracefully, and reschedules the pod on a replacement node. In-flight requests fail fast and are retried by callers; interruption never produces partial results.
Scale-to-Zero (On-Demand Mode)
Set Auto-shutdown (or auto_shutdown_minutes via REST) to scale the model
to zero replicas after N idle minutes. The first request after idle wakes the
model back to its configured replica count. For scheduled batch workloads a
short timeout (e.g. 10 minutes) trims the idle GPU tail without affecting the
run -- model pods are continuously busy while a workflow is active.
Model Operations
Start and Stop
Start and stop a deployed model from its details page without deleting it. Stopping frees its compute and keeps its configuration; starting brings it back with the same settings (and is checked against its budget).
Changing a Deployed Model
Edit a running model on its Settings tab and click Save. The model is changed in place; it is never deleted or redeployed, and its endpoint stays the same:
- CPU, memory, disk and GPU: the model restarts at exactly the new size. A bigger size is checked against the model's budget first; if the budget refuses it, nothing changes.
- Replicas, scheduling and scaling: take effect on the running model without a restart. The combination must be valid (an on-demand model keeps an idle window; auto-scaling keeps its three settings), or Save shows why and nothing changes.
Over REST: PATCH /api/v1/ai-gateway/models/:id/resources resizes, and PATCH /api/v1/ai-gateway/models/:id changes replicas, scheduling and scaling. See AI Models API.
Status and Logs
The details page shows the model's status, replicas and resource use, and its Logs tab shows the image build, deployment events and runtime logs. Over REST: GET /api/v1/ai-gateway/models/:id/status and GET /api/v1/ai-gateway/models/:id/logs.
Customizing Dockerfile
For advanced use cases, customize the Docker image for your model deployment. The platform provides Dockerfile templates as starting points:
# Example: Custom vLLM Dockerfile (based on platform template)
FROM nvcr.io/nvidia/pytorch:23.10-py3
# Install vLLM and dependencies
RUN pip install vllm transformers accelerate
# Set environment variables
ENV MODEL_NAME=mistralai/Mistral-7B-Instruct-v0.2
ENV PORT=8000
ENV GPU_MEMORY_UTILIZATION=0.95
ENV MAX_MODEL_LEN=8192
# Create model directory
RUN mkdir -p /models
WORKDIR /models
# Expose port
EXPOSE 8000
# Health check endpoint
HEALTHCHECK CMD curl -f http://localhost:8000/health || exit 1
# Start vLLM server
CMD python -m vllm.entrypoints.openai.api_server \
--model $MODEL_NAME \
--port $PORT \
--gpu-memory-utilization $GPU_MEMORY_UTILIZATION \
--max-model-len $MAX_MODEL_LEN \
--trust-remote-code
Dockerfile Templates
The deploy form starts from the platform's pre-built model templates (over REST: GET /api/v1/ai-gateway/models/prebuilt), for example Mistral 7B and Llama models on vLLM, and custom vLLM models.
Common customizations:
- Additional Python packages
- Custom tokenizers
- Preprocessing scripts
- Model quantization
- Multi-GPU configuration
Deployment API Endpoints
| Endpoint | Method | Description |
|---|---|---|
/api/v1/ai-gateway/models/prebuilt | GET | Pre-built models you can deploy |
/api/v1/ai-gateway/models | POST | Register a self-hosted model |
/api/v1/ai-gateway/models/:id/deploy | POST | Deploy it (scheduling, scaling, instance, spot) |
/api/v1/ai-gateway/models/:id | PUT | Change replicas, scheduling and scaling in place |
/api/v1/ai-gateway/models/:id/resources | PATCH | Resize CPU, memory, disk, GPU in place |
/api/v1/ai-gateway/models/:id/start | POST | Start a stopped model |
/api/v1/ai-gateway/models/:id/stop | POST | Stop a running model (keeps its configuration) |
/api/v1/ai-gateway/models/:id/status | GET | Deployment status |
/api/v1/ai-gateway/models/:id/logs | GET | Build, deployment and runtime logs |
Full reference: AI Models API.
Next Steps
- Fine-tune a model on your custom data
- Configure autoscaling to handle varying inference loads
- Set up monitoring to track model performance
- Optimize costs with right-sized resources