Self-Hosted Models
Deploy 169 pre-configured open-source AI models on your own infrastructure. Each model comes with a production-ready Dockerfile and is routed through the AI Gateway's standardized API endpoints -- the same endpoints used by third-party providers like OpenAI, Anthropic, and Google.
Model Catalogue
| Category | Count | Model Types |
|---|---|---|
| LLM | 83 | Chat, code generation, reasoning |
| Multimodal | 32 | Vision-language models (VLMs) |
| Embedding | 12 | Text embeddings for RAG/search |
| Audio | 34 | Text-to-speech (16), speech-to-text (18) |
| Video | 6 | Text-to-video generation |
| Image | 1 | Image generation (Stable Diffusion) |
| NLP | 1 | Specialized NLP tasks |
Provider & Engine Architecture
All self-hosted models use provider: self_hosted. The engine field determines how the AI Gateway routes requests -- one provider maps to many engines.
Provider vs Engine
| Concept | What It Means | Examples |
|---|---|---|
| Provider | Who provides the model (the vendor) | self_hosted, openai, anthropic, custom |
| Engine | How to communicate with the inference server | vllm, transformers, whisper, tts-engine, video-engine, custom |
Engine Routing
| Engine | Internal Handler | Used For | API Format |
|---|---|---|---|
vllm | VLLMProvider | LLM chat, multimodal, embedding | OpenAI-compatible /v1/chat/completions |
transformers | CustomProvider | Custom transformers servers | Configurable endpoints |
custom | CustomProvider | Generic custom servers | Configurable endpoints |
whisper | CustomProvider | Speech-to-text models | /v1/audio/transcriptions |
tts-engine | CustomProvider | Text-to-speech models | /v1/audio/speech |
video-engine | CustomProvider | Video generation models | /v1/video/generations |
sentence-transformers | SentenceTransformersProvider | The platform embedding model, run inside the AI Gateway | /v1/embeddings |
How Routing Works
User Request (e.g., POST /v1/chat/completions)
↓
AI Gateway receives request with model ID
↓
Looks up model config: provider=self_hosted, engine=vllm
↓
Provider Factory sees provider=self_hosted, reads engine field:
- engine: vllm → VLLMProvider
- engine: transformers/custom/whisper/tts-engine/video-engine → CustomProvider
- engine: sentence-transformers → SentenceTransformersProvider (in the gateway)
↓
Provider forwards request to self-hosted model's internal endpoint
↓
Response is normalized to standard format and returned
Engine Capabilities
VLLMProvider (engines: vllm, transformers):
| Capability | Supported | Endpoint |
|---|---|---|
| Chat completions | Yes | /v1/chat/completions |
| Streaming | Yes | /v1/chat/completions (SSE) |
| Embeddings | Yes | /v1/embeddings |
| Image generation | Yes | /v1/images/generations |
| Vision/Multimodal | Yes | /v1/chat/completions with image content |
| Tool calling | Models whose catalog entry enables it (see Tool Calling) | /v1/chat/completions with tools |
CustomProvider (engines: custom, whisper, tts-engine, video-engine):
| Capability | Supported | Endpoint |
|---|---|---|
| Chat completions | Yes | /v1/chat/completions |
| Streaming | Yes | /v1/chat/completions (SSE) |
| Embeddings | Yes | /v1/embeddings |
| Text-to-speech | Yes | /v1/audio/speech |
| Speech-to-text | Yes | /v1/audio/transcriptions |
| Image generation | Yes | /v1/images/generations |
| Video generation | Yes | /v1/video/generations |
Tool Calling
A self-hosted model calls tools only when its catalog entry enables tool calling: the catalog starts those models' servers with tool calling on, for the models whose chat format supports it. Their model cards show a Tool calling badge (hover it for the tool-call parser). The catalog enables it on the Qwen 2.5, Qwen 2.5 Coder, Qwen3, QwQ, Qwen3-VL and Qwen3.5 models, Llama 3.1, 3.2 (including 3.2 Vision) and 3.3, Mistral 7B v0.3 and Mistral Nemo, Gemma 3, InternLM2, Voxtral Mini, the MiniMax M2 models and GLM-5. Models whose chat format has no tool calls (for example Llama 2, Phi-3, Gemma 2 and Qwen2.5-VL) are served without it.
Every model that calls tools is served with a context of at least 32,768 tokens, or its full native context where the GPU allows it, so an agent's instructions, tools and conversation fit. A deployed model records its context window, and an agent on it sizes its context policy from that.
A deployed model records whether it can call tools, and the AI Gateway goes by that record:
- A model that can call tools receives a request's tools, its tool choice and whether to call tools in parallel, and returns its tool calls, streamed or not.
- A model that cannot call tools is refused a request that includes tools, with the reason. It is never answered as if it had no tools.
- A model deployed before its catalog entry enabled tool calling keeps the version it was deployed with. Redeploy it to take the catalog's current version.
Web search
Models with web search of their own use it. For any other model, the AI Gateway runs the search itself: it offers the model a search tool, runs the searches the model asks for, and gives it the results. That needs tool calling, so on a self-hosted model web search works only when the model can call tools. A request for web search on a model that can neither call tools nor search is refused with the reason; nothing is searched and the model does not answer as if it had searched.
Inference Engines
vLLM Engine (113 models)
The majority of models use the vLLM inference engine -- a high-performance serving framework with PagedAttention, continuous batching, and OpenAI-compatible API.
Base Image: vllm/vllm-openai:v0.16.0
Key vLLM flags:
--tensor-parallel-size N-- Distribute model across N GPUs--gpu-memory-utilization 0.9-- GPU memory fraction to use--max-model-len N-- Maximum context length--served-model-name name-- Model name exposed via API--trust-remote-code-- Required for some HuggingFace models--enable-auto-tool-choice --tool-call-parser hermes-- Enable tool calling (both are needed; the parser matches the model's tool-call format)--enable-reasoning-- Enable chain-of-thought reasoning--limit-mm-per-prompt image=4-- Limit multimodal inputs per request
Example: Llama 3.1 8B (one L40S on g6e.xlarge, its full 131,072-token context, tool calling on)
FROM vllm/vllm-openai:v0.16.0
ENV MODEL_NAME="meta-llama/Llama-3.1-8B-Instruct"
ENTRYPOINT ["python3", "-m", "vllm.entrypoints.openai.api_server"]
CMD ["--model", "meta-llama/Llama-3.1-8B-Instruct",
"--host", "0.0.0.0", "--port", "8000",
"--tensor-parallel-size", "1",
"--gpu-memory-utilization", "0.90",
"--max-model-len", "131072",
"--max-num-seqs", "5",
"--enable-auto-tool-choice", "--tool-call-parser", "llama3_json",
"--chat-template", "/vllm-workspace/examples/tool_chat_template_llama3.1_json.jinja"]
Transformers Engine (24 models)
Custom Python servers using HuggingFace Transformers directly. Used when a model isn't supported by vLLM or needs custom preprocessing.
Must expose: Standard OpenAI-compatible /v1/chat/completions endpoint accepting JSON:
{
"model": "model-name",
"messages": [{"role": "user", "content": "Hello"}],
"max_tokens": 512
}
For multimodal transformers models, the server must handle both text-only and image+text messages in OpenAI format:
{
"messages": [{
"role": "user",
"content": [
{"type": "text", "text": "What do you see?"},
{"type": "image_url", "image_url": {"url": "data:image/png;base64,..."}}
]
}]
}
TTS Engine (16 models)
Text-to-speech models that expose /v1/audio/speech endpoint.
Request format:
{
"input": "Text to speak",
"voice": "default",
"model": "model-name",
"response_format": "mp3",
"speed": 1.0
}
Response: Binary audio stream.
Whisper Engine (5 models)
Speech-to-text models using OpenAI Whisper-compatible API at /v1/audio/transcriptions.
Request format: multipart/form-data with audio file upload.
Response:
{
"text": "Transcribed text content"
}
Video Engine (6 models)
Text-to-video generation models that expose /v1/video/generations endpoint.
Models: CogVideoX-2B, CogVideoX-5B, HunyuanVideo, Wan2.1-T2V-1.3B, Wan2.1-T2V-14B, LTX-Video
Request format:
{
"prompt": "A ball bouncing in slow motion",
"duration": 5,
"resolution": "1080p",
"aspect_ratio": "16:9",
"fps": 24
}
Response: Either JSON with job ID (async) or binary video data (sync), depending on the model.
Model Types & API Endpoints
Each model type maps to a specific AI Gateway endpoint:
| Model Type | Gateway Endpoint | Proxy Path |
|---|---|---|
chat | POST /v1/chat/completions | /api/v1/ai-gateway/chat/completions |
multimodal | POST /v1/chat/completions | /api/v1/ai-gateway/chat/completions |
embedding | POST /v1/embeddings | /api/v1/ai-gateway/embeddings |
text_to_speech | POST /v1/audio/speech | /api/v1/ai-gateway/audio/speech |
speech_to_text | POST /v1/audio/transcriptions | /api/v1/ai-gateway/audio/transcriptions |
image_generation | POST /v1/generations/images | /api/v1/ai-gateway/generations/images |
video_generation | POST /v1/generations/videos | /api/v1/ai-gateway/generations/videos |
Important: Multimodal models use the same /v1/chat/completions endpoint as chat models. The gateway accepts both chat and multimodal model types for chat completions.
The platform embedding model
Every platform ships with Snowflake Arctic Embed S, a small embedding model (384 dimensions) the AI Gateway runs inside its own process (engine: sentence-transformers). It belongs to the platform, so every user may use it; no one shares, reassigns or edits it, and its Permissions tab says so.
It has no deployment of its own. Its page shows Runs in: The AI Gateway in place of a deployment ID, and no CPU, memory, replicas, disk, Metrics or Logs, because it shares the gateway's. Its request figures count your own requests, as every usage figure does (see Monitoring).
Deployment Lifecycle
1. Create Deployment
When you deploy a self-hosted model, the platform:
- Reads the model's Dockerfile from the catalogue
- Builds the Docker image using a K8s build job (Docker-in-Docker)
- Pushes the image to ECR
- Creates a Kubernetes Deployment + Service in the org namespace
- Registers the model in the AI Gateway with
provider: self_hostedand the model'senginefield
2. Model Status Flow
building → deploying → active → running
→ error (build/deploy failure)
- building: Docker image is being built
- deploying: K8s resources are being created
- active: K8s Deployment exists, pod is starting
- running: Model is serving inference requests
- error: Build or deployment failed
3. Inference Readiness
After a model reaches "active" status, the inference server still needs time to:
- Pull the Docker image
- Download model weights from HuggingFace
- Load weights into GPU memory
- Start the HTTP server
The gateway handles this transparently -- requests return 503 until the model is ready.
GPU Requirements
Each catalog model runs on the cheapest GPU that holds its weights and its context (at least 32,768 tokens where the model supports it), on one GPU wherever one is enough. Its template names that GPU type and node, and its CPU and memory fill the node, so a deployment lands on exactly that node. The deploy form starts on the template's GPU type rather than Auto; with Auto, a model could land on a smaller GPU than it was sized for.
| GPUs | Node | Example models |
|---|---|---|
| None (CPU) | t3.medium, t3.large | Embeddings, Whisper tiny and small, smaller TTS models |
| 1x T4 16GB | g4dn.xlarge | Qwen 2.5 0.5B and 3B, Llama 3.2 1B and 3B, Phi-2, Phi-3 mini 4K, most speech and TTS models |
| 1x L4 24GB | g6.xlarge, g6.2xlarge | 7B-8B LLMs at up to 32K (Qwen 2.5 7B, Mistral 7B), Gemma 3 4B, Voxtral Mini |
| 1x A10G 24GB | g5.xlarge to g5.4xlarge | Image, video and speech models built for it, MiniCPM-V 2.6 |
| 1x L40S 48GB | g6e.xlarge, g6e.2xlarge | 8B-14B LLMs at long context (Llama 3.1 8B, Qwen3 8B and 14B), Mistral Nemo, Gemma 3 12B, 7B-13B vision models |
| 1x H100 80GB | p5.4xlarge | 20B-35B LLMs (Qwen 2.5 and Qwen3 32B, QwQ 32B, Qwen3 30B-A3B), Gemma 3 27B, Qwen3.5 27B and 35B-A3B |
| 4x L40S 48GB | g6e.12xlarge | 70B-72B LLMs, Mixtral 8x7B, Command R |
| 8x A100 40GB | p4d.24xlarge | Llama 3.2 90B Vision, Qwen2.5-VL 72B, InternVL3 78B, Qwen3.5 122B-A10B, MiniMax M2 |
| 8x A100 80GB | p4de.24xlarge | Llama 3.1 405B FP8, Qwen3.5 397B-A17B (FP8 checkpoint) |
| 8x H200 141GB | p5en.48xlarge | GLM-5 FP8 |
The AI pool provisions nodes up to 4xlarge unless Compute settings allow larger ones, so the multi-GPU models above need their node size allowed in Compute before they deploy.
Adding Custom Models
To add a new self-hosted model to the catalogue, add an entry to self-hosted-models.js:
{
id: 'my-custom-model',
name: 'My Custom Model',
description: 'Description of the model',
category: 'llm', // llm, multimodal, embedding, audio, video, image
provider: 'vendor-name',
modelType: 'chat', // chat, multimodal, embedding, text_to_speech, etc.
engine: 'vllm', // vllm, transformers, custom, whisper, tts-engine, video-engine
dockerfile: `...`, // Full Dockerfile content
defaultPort: 8000,
defaultResources: { cpu: '3620m', memory: '13Gi', gpu: 1, gpu_type: 'nvidia-l4', disk: '50Gi' },
recommendedInstance: 'g6.xlarge', // its CPU and memory must fit what this node gives a pod (the whole g6.xlarge here)
tags: ['llm', 'chat', '7b'],
modelSize: '14GB',
documentation: '# My Model\n\nModel documentation here.'
}
Engine Selection Guide
| Choose This Engine | When |
|---|---|
vllm | Model is supported by vLLM (most LLMs, VLMs, embeddings) |
transformers | Custom HuggingFace model not in vLLM, needs custom preprocessing |
custom | Non-standard inference server with custom API |
whisper | OpenAI Whisper-compatible STT model |
tts-engine | Text-to-speech model exposing /v1/audio/speech |
video-engine | Video generation model exposing /v1/video/generations |
Server Requirements
All self-hosted model servers must expose:
- Health endpoint:
GET /healthreturning{"status": "ok"} - Inference endpoint: The appropriate endpoint for the model type (see table above)
- Port 8000: Default port (configurable via
defaultPort)
The inference endpoint must accept and return standard JSON -- not multipart form data or custom formats (except STT which uses multipart for file upload).