Skip to main content

Self-Hosted Models

Deploy 169 pre-configured open-source AI models on your own infrastructure. Each model comes with a production-ready Dockerfile and is routed through the AI Gateway's standardized API endpoints -- the same endpoints used by third-party providers like OpenAI, Anthropic, and Google.

Model Catalogue​

CategoryCountModel Types
LLM83Chat, code generation, reasoning
Multimodal32Vision-language models (VLMs)
Embedding12Text embeddings for RAG/search
Audio34Text-to-speech (16), speech-to-text (18)
Video6Text-to-video generation
Image1Image generation (Stable Diffusion)
NLP1Specialized NLP tasks

Provider & Engine Architecture​

All self-hosted models use provider: self_hosted. The engine field determines how the AI Gateway routes requests -- one provider maps to many engines.

Provider vs Engine​

ConceptWhat It MeansExamples
ProviderWho provides the model (the vendor)self_hosted, openai, anthropic, custom
EngineHow to communicate with the inference servervllm, transformers, whisper, tts-engine, video-engine, custom

Engine Routing​

EngineInternal HandlerUsed ForAPI Format
vllmVLLMProviderLLM chat, multimodal, embeddingOpenAI-compatible /v1/chat/completions
transformersCustomProviderCustom transformers serversConfigurable endpoints
customCustomProviderGeneric custom serversConfigurable endpoints
whisperCustomProviderSpeech-to-text models/v1/audio/transcriptions
tts-engineCustomProviderText-to-speech models/v1/audio/speech
video-engineCustomProviderVideo generation models/v1/video/generations
sentence-transformersSentenceTransformersProviderThe platform embedding model, run inside the AI Gateway/v1/embeddings

How Routing Works​

User Request (e.g., POST /v1/chat/completions)
↓
AI Gateway receives request with model ID
↓
Looks up model config: provider=self_hosted, engine=vllm
↓
Provider Factory sees provider=self_hosted, reads engine field:
- engine: vllm → VLLMProvider
- engine: transformers/custom/whisper/tts-engine/video-engine → CustomProvider
- engine: sentence-transformers → SentenceTransformersProvider (in the gateway)
↓
Provider forwards request to self-hosted model's internal endpoint
↓
Response is normalized to standard format and returned

Engine Capabilities​

VLLMProvider (engines: vllm, transformers):

CapabilitySupportedEndpoint
Chat completionsYes/v1/chat/completions
StreamingYes/v1/chat/completions (SSE)
EmbeddingsYes/v1/embeddings
Image generationYes/v1/images/generations
Vision/MultimodalYes/v1/chat/completions with image content
Tool callingModels whose catalog entry enables it (see Tool Calling)/v1/chat/completions with tools

CustomProvider (engines: custom, whisper, tts-engine, video-engine):

CapabilitySupportedEndpoint
Chat completionsYes/v1/chat/completions
StreamingYes/v1/chat/completions (SSE)
EmbeddingsYes/v1/embeddings
Text-to-speechYes/v1/audio/speech
Speech-to-textYes/v1/audio/transcriptions
Image generationYes/v1/images/generations
Video generationYes/v1/video/generations

Tool Calling​

A self-hosted model calls tools only when its catalog entry enables tool calling: the catalog starts those models' servers with tool calling on, for the models whose chat format supports it. Their model cards show a Tool calling badge (hover it for the tool-call parser). The catalog enables it on the Qwen 2.5, Qwen 2.5 Coder, Qwen3, QwQ, Qwen3-VL and Qwen3.5 models, Llama 3.1, 3.2 (including 3.2 Vision) and 3.3, Mistral 7B v0.3 and Mistral Nemo, Gemma 3, InternLM2, Voxtral Mini, the MiniMax M2 models and GLM-5. Models whose chat format has no tool calls (for example Llama 2, Phi-3, Gemma 2 and Qwen2.5-VL) are served without it.

Every model that calls tools is served with a context of at least 32,768 tokens, or its full native context where the GPU allows it, so an agent's instructions, tools and conversation fit. A deployed model records its context window, and an agent on it sizes its context policy from that.

A deployed model records whether it can call tools, and the AI Gateway goes by that record:

  • A model that can call tools receives a request's tools, its tool choice and whether to call tools in parallel, and returns its tool calls, streamed or not.
  • A model that cannot call tools is refused a request that includes tools, with the reason. It is never answered as if it had no tools.
  • A model deployed before its catalog entry enabled tool calling keeps the version it was deployed with. Redeploy it to take the catalog's current version.

Models with web search of their own use it. For any other model, the AI Gateway runs the search itself: it offers the model a search tool, runs the searches the model asks for, and gives it the results. That needs tool calling, so on a self-hosted model web search works only when the model can call tools. A request for web search on a model that can neither call tools nor search is refused with the reason; nothing is searched and the model does not answer as if it had searched.

Inference Engines​

vLLM Engine (113 models)​

The majority of models use the vLLM inference engine -- a high-performance serving framework with PagedAttention, continuous batching, and OpenAI-compatible API.

Base Image: vllm/vllm-openai:v0.16.0

Key vLLM flags:

  • --tensor-parallel-size N -- Distribute model across N GPUs
  • --gpu-memory-utilization 0.9 -- GPU memory fraction to use
  • --max-model-len N -- Maximum context length
  • --served-model-name name -- Model name exposed via API
  • --trust-remote-code -- Required for some HuggingFace models
  • --enable-auto-tool-choice --tool-call-parser hermes -- Enable tool calling (both are needed; the parser matches the model's tool-call format)
  • --enable-reasoning -- Enable chain-of-thought reasoning
  • --limit-mm-per-prompt image=4 -- Limit multimodal inputs per request

Example: Llama 3.1 8B (one L40S on g6e.xlarge, its full 131,072-token context, tool calling on)

FROM vllm/vllm-openai:v0.16.0
ENV MODEL_NAME="meta-llama/Llama-3.1-8B-Instruct"
ENTRYPOINT ["python3", "-m", "vllm.entrypoints.openai.api_server"]
CMD ["--model", "meta-llama/Llama-3.1-8B-Instruct",
"--host", "0.0.0.0", "--port", "8000",
"--tensor-parallel-size", "1",
"--gpu-memory-utilization", "0.90",
"--max-model-len", "131072",
"--max-num-seqs", "5",
"--enable-auto-tool-choice", "--tool-call-parser", "llama3_json",
"--chat-template", "/vllm-workspace/examples/tool_chat_template_llama3.1_json.jinja"]

Transformers Engine (24 models)​

Custom Python servers using HuggingFace Transformers directly. Used when a model isn't supported by vLLM or needs custom preprocessing.

Must expose: Standard OpenAI-compatible /v1/chat/completions endpoint accepting JSON:

{
"model": "model-name",
"messages": [{"role": "user", "content": "Hello"}],
"max_tokens": 512
}

For multimodal transformers models, the server must handle both text-only and image+text messages in OpenAI format:

{
"messages": [{
"role": "user",
"content": [
{"type": "text", "text": "What do you see?"},
{"type": "image_url", "image_url": {"url": "data:image/png;base64,..."}}
]
}]
}

TTS Engine (16 models)​

Text-to-speech models that expose /v1/audio/speech endpoint.

Request format:

{
"input": "Text to speak",
"voice": "default",
"model": "model-name",
"response_format": "mp3",
"speed": 1.0
}

Response: Binary audio stream.

Whisper Engine (5 models)​

Speech-to-text models using OpenAI Whisper-compatible API at /v1/audio/transcriptions.

Request format: multipart/form-data with audio file upload.

Response:

{
"text": "Transcribed text content"
}

Video Engine (6 models)​

Text-to-video generation models that expose /v1/video/generations endpoint.

Models: CogVideoX-2B, CogVideoX-5B, HunyuanVideo, Wan2.1-T2V-1.3B, Wan2.1-T2V-14B, LTX-Video

Request format:

{
"prompt": "A ball bouncing in slow motion",
"duration": 5,
"resolution": "1080p",
"aspect_ratio": "16:9",
"fps": 24
}

Response: Either JSON with job ID (async) or binary video data (sync), depending on the model.


Model Types & API Endpoints​

Each model type maps to a specific AI Gateway endpoint:

Model TypeGateway EndpointProxy Path
chatPOST /v1/chat/completions/api/v1/ai-gateway/chat/completions
multimodalPOST /v1/chat/completions/api/v1/ai-gateway/chat/completions
embeddingPOST /v1/embeddings/api/v1/ai-gateway/embeddings
text_to_speechPOST /v1/audio/speech/api/v1/ai-gateway/audio/speech
speech_to_textPOST /v1/audio/transcriptions/api/v1/ai-gateway/audio/transcriptions
image_generationPOST /v1/generations/images/api/v1/ai-gateway/generations/images
video_generationPOST /v1/generations/videos/api/v1/ai-gateway/generations/videos

Important: Multimodal models use the same /v1/chat/completions endpoint as chat models. The gateway accepts both chat and multimodal model types for chat completions.

The platform embedding model​

Every platform ships with Snowflake Arctic Embed S, a small embedding model (384 dimensions) the AI Gateway runs inside its own process (engine: sentence-transformers). It belongs to the platform, so every user may use it; no one shares, reassigns or edits it, and its Permissions tab says so.

It has no deployment of its own. Its page shows Runs in: The AI Gateway in place of a deployment ID, and no CPU, memory, replicas, disk, Metrics or Logs, because it shares the gateway's. Its request figures count your own requests, as every usage figure does (see Monitoring).


Deployment Lifecycle​

1. Create Deployment​

When you deploy a self-hosted model, the platform:

  1. Reads the model's Dockerfile from the catalogue
  2. Builds the Docker image using a K8s build job (Docker-in-Docker)
  3. Pushes the image to ECR
  4. Creates a Kubernetes Deployment + Service in the org namespace
  5. Registers the model in the AI Gateway with provider: self_hosted and the model's engine field

2. Model Status Flow​

building → deploying → active → running
→ error (build/deploy failure)
  • building: Docker image is being built
  • deploying: K8s resources are being created
  • active: K8s Deployment exists, pod is starting
  • running: Model is serving inference requests
  • error: Build or deployment failed

3. Inference Readiness​

After a model reaches "active" status, the inference server still needs time to:

  • Pull the Docker image
  • Download model weights from HuggingFace
  • Load weights into GPU memory
  • Start the HTTP server

The gateway handles this transparently -- requests return 503 until the model is ready.


GPU Requirements​

Each catalog model runs on the cheapest GPU that holds its weights and its context (at least 32,768 tokens where the model supports it), on one GPU wherever one is enough. Its template names that GPU type and node, and its CPU and memory fill the node, so a deployment lands on exactly that node. The deploy form starts on the template's GPU type rather than Auto; with Auto, a model could land on a smaller GPU than it was sized for.

GPUsNodeExample models
None (CPU)t3.medium, t3.largeEmbeddings, Whisper tiny and small, smaller TTS models
1x T4 16GBg4dn.xlargeQwen 2.5 0.5B and 3B, Llama 3.2 1B and 3B, Phi-2, Phi-3 mini 4K, most speech and TTS models
1x L4 24GBg6.xlarge, g6.2xlarge7B-8B LLMs at up to 32K (Qwen 2.5 7B, Mistral 7B), Gemma 3 4B, Voxtral Mini
1x A10G 24GBg5.xlarge to g5.4xlargeImage, video and speech models built for it, MiniCPM-V 2.6
1x L40S 48GBg6e.xlarge, g6e.2xlarge8B-14B LLMs at long context (Llama 3.1 8B, Qwen3 8B and 14B), Mistral Nemo, Gemma 3 12B, 7B-13B vision models
1x H100 80GBp5.4xlarge20B-35B LLMs (Qwen 2.5 and Qwen3 32B, QwQ 32B, Qwen3 30B-A3B), Gemma 3 27B, Qwen3.5 27B and 35B-A3B
4x L40S 48GBg6e.12xlarge70B-72B LLMs, Mixtral 8x7B, Command R
8x A100 40GBp4d.24xlargeLlama 3.2 90B Vision, Qwen2.5-VL 72B, InternVL3 78B, Qwen3.5 122B-A10B, MiniMax M2
8x A100 80GBp4de.24xlargeLlama 3.1 405B FP8, Qwen3.5 397B-A17B (FP8 checkpoint)
8x H200 141GBp5en.48xlargeGLM-5 FP8

The AI pool provisions nodes up to 4xlarge unless Compute settings allow larger ones, so the multi-GPU models above need their node size allowed in Compute before they deploy.


Adding Custom Models​

To add a new self-hosted model to the catalogue, add an entry to self-hosted-models.js:

{
id: 'my-custom-model',
name: 'My Custom Model',
description: 'Description of the model',
category: 'llm', // llm, multimodal, embedding, audio, video, image
provider: 'vendor-name',
modelType: 'chat', // chat, multimodal, embedding, text_to_speech, etc.
engine: 'vllm', // vllm, transformers, custom, whisper, tts-engine, video-engine
dockerfile: `...`, // Full Dockerfile content
defaultPort: 8000,
defaultResources: { cpu: '3620m', memory: '13Gi', gpu: 1, gpu_type: 'nvidia-l4', disk: '50Gi' },
recommendedInstance: 'g6.xlarge', // its CPU and memory must fit what this node gives a pod (the whole g6.xlarge here)
tags: ['llm', 'chat', '7b'],
modelSize: '14GB',
documentation: '# My Model\n\nModel documentation here.'
}

Engine Selection Guide​

Choose This EngineWhen
vllmModel is supported by vLLM (most LLMs, VLMs, embeddings)
transformersCustom HuggingFace model not in vLLM, needs custom preprocessing
customNon-standard inference server with custom API
whisperOpenAI Whisper-compatible STT model
tts-engineText-to-speech model exposing /v1/audio/speech
video-engineVideo generation model exposing /v1/video/generations

Server Requirements​

All self-hosted model servers must expose:

  1. Health endpoint: GET /health returning {"status": "ok"}
  2. Inference endpoint: The appropriate endpoint for the model type (see table above)
  3. Port 8000: Default port (configurable via defaultPort)

The inference endpoint must accept and return standard JSON -- not multipart form data or custom formats (except STT which uses multipart for file upload).