Providers
The AI Gateway supports 15 third-party AI providers with 268 certified models, plus self-hosted models. All providers use a unified API interface. The models listed here are each vendor's current models (October 2026); models a vendor retires leave the catalog (see Certified Models).
Provider Summary
| Provider | Models | Capabilities |
|---|---|---|
| OpenAI | 109 | Chat, Embeddings, TTS, STT, Realtime, Images |
| Anthropic | 14 | Chat with Vision |
| Google (Gemini) | 35 | Chat, Embeddings, TTS, Images, Video |
| Mistral | 33 | Chat, Embeddings, STT, OCR, Moderation |
| Cohere | 25 | Chat, Embeddings, Reranking |
| Grok (xAI) | 11 | Chat, Images |
| DeepSeek | 2 | Chat, Reasoning |
| ElevenLabs | 13 | TTS, Speech-to-Speech, Music, Sound Effects |
| Mubert | - | Music Generation |
| Deepgram | - | Speech-to-Text |
| Stability AI | 8 | Image Generation |
| Black Forest Labs | 11 | Image Generation (FLUX) |
| Runway | 5 | Video Generation |
| Luma AI | 2 | Video Generation |
| Custom Provider | - | Any OpenAI-compatible API |
| Self-Hosted | 169 | Self-hosted models (engine determines protocol: vllm, transformers, whisper, tts-engine, video-engine, custom) |
Provider API Keys and Adding a Model
Save a provider's API key under AI Gateway > API Keys, choosing the provider from the same cards as the Third Party AI page. Test checks a saved key by listing the provider's models with it; no model is called, so the test does not depend on any one model being available.
On the Third Party AI page, pick a provider card and choose the model from the provider's certified models (the field shows a current model of that provider as its example, such as gpt-6.1-sol for OpenAI or claude-opus-5-5 for Anthropic). Test Connection sends that model a minimal request with your key.
OpenAI
Models for chat, reasoning, embeddings, audio, realtime voice, and image generation.
Capabilities: Chat, Embeddings, Text-to-Speech, Speech-to-Text, Realtime, Image Generation
Featured Models
| Model | Type | Best For |
|---|---|---|
openai/gpt-6-astra | Chat | OpenAI's most capable model, for the most demanding work |
openai/gpt-6.1-sol | Chat | Complex coding and professional tasks at lower cost than Astra |
openai/gpt-6-luna | Chat | Efficient, high-volume tasks |
openai/text-embedding-3-large | Embedding | High-quality embeddings |
openai/gpt-4o-mini-tts | TTS | Speech synthesis |
openai/gpt-transcribe | STT | Audio transcription |
openai/gpt-realtime-2.1 | Realtime | Voice agents with reasoning and tools |
openai/gpt-image-2.5-sunburst | Image | Highest-quality image generation |
openai/gpt-image-2.5-flare | Image | Fast everyday image generation |
Anthropic
Claude models with strong reasoning, coding, and long-context capabilities (1M-token context window).
Capabilities: Chat with Vision, Function Calling, Web Search
Featured Models
| Model | Type | Best For |
|---|---|---|
anthropic/claude-opus-5-5 | Chat | Long-running agentic coding and knowledge work |
anthropic/claude-fable-5-1 | Chat | Demanding reasoning and long-horizon agentic work |
anthropic/claude-sonnet-5-5 | Chat | Best combination of speed and intelligence |
anthropic/claude-haiku-5-5 | Chat | High-volume, latency-sensitive tasks |
Google (Gemini)
Multimodal models with large context windows and diverse capabilities.
Capabilities: Chat, Embeddings, TTS, Image Generation, Video Generation
Featured Models
| Model | Type | Best For |
|---|---|---|
gemini/gemini-3.8-flash | Chat | Fast multimodal processing with search grounding |
gemini/gemini-3.1-pro-preview | Chat | Most capable Gemini model |
gemini/gemini-3.5-flash-lite | Chat | Fast, cost-effective |
gemini/gemini-embedding-2 | Embedding | Text embeddings |
gemini/gemini-3.8-flash-tts | TTS | Speech synthesis |
gemini/gemini-nano-banana-2.1 | Image | Image generation and editing |
gemini/veo-3.1-generate-preview | Video | Video generation |
Note: In the backend provider enum, Google models use the gemini provider identifier. When configuring models via the API, use gemini as the provider value.
Mistral
European AI with strong multilingual and code capabilities.
Capabilities: Chat, Embeddings, Speech-to-Text, OCR, Moderation
Featured Models
| Model | Type | Best For |
|---|---|---|
mistral/mistral-large-4 | Chat | Most capable Mistral model |
mistral/mistral-medium-2604 | Chat | Mistral Medium 3.5, balanced quality and cost |
mistral/mistral-small-2603 | Chat | Mistral Small 4, fast and cost-effective |
mistral/codestral-latest | Chat | Code generation |
mistral/mistral-embed | Embedding | Text embeddings |
mistral/mistral-ocr-4-1 | OCR | Document text and layout extraction |
Cohere
Enterprise-focused models with RAG and reranking specialization.
Capabilities: Chat, Embeddings, Reranking
Featured Models
| Model | Type | Best For |
|---|---|---|
cohere/command-a-plus-05-2026 | Chat | Cohere's most capable model |
cohere/command-a-reasoning-08-2025 | Chat | Reasoning tasks |
cohere/command-a-vision-07-2025 | Chat | Vision tasks |
cohere/embed-v5.0-pro | Embedding | Multilingual, multimodal embeddings |
cohere/embed-v5.0-fast | Embedding | Faster multimodal embeddings |
cohere/rerank-v4.0-pro | Rerank | Search result reranking |
Grok (xAI)
xAI's models with real-time knowledge and image generation.
Capabilities: Chat, Vision, Image Generation
Featured Models
| Model | Type | Best For |
|---|---|---|
grok/grok-4.7 | Chat | Most capable Grok model, code and chat |
grok/grok-4.3 | Chat | 1M-token context window |
grok/grok-4.20-0309-non-reasoning | Chat | Fast answers without reasoning |
grok/grok-build-0.1 | Chat | Coding |
grok/grok-imagine-image-quality | Image | Image generation |
DeepSeek
Strong reasoning and coding capabilities. Thinking is on when a request sets a reasoning effort.
Capabilities: Chat, Reasoning
Featured Models
| Model | Type | Best For |
|---|---|---|
deepseek/deepseek-v4-pro | Chat | Complex reasoning and coding |
deepseek/deepseek-flash | Chat | Fast, cost-effective chat |
ElevenLabs
Voice synthesis and audio generation.
Capabilities: Text-to-Speech, Speech-to-Speech, Music Generation, Sound Effects
Featured Models
| Model | Type | Best For |
|---|---|---|
elevenlabs/eleven_v4 | TTS | Most expressive speech, with audio tags |
elevenlabs/eleven_v4_turbo | TTS | Fast expressive speech |
elevenlabs/eleven_multilingual_v2 | TTS | Multilingual voice synthesis |
elevenlabs/eleven_flash_v2_5 | TTS | Ultra-fast TTS |
elevenlabs/eleven_music | Music | Music generation |
elevenlabs/sound-effects-v2 | Audio | Sound effects |
Mubert
AI-powered music generation platform.
Capabilities: Music Generation
Featured Models
| Model | Type | Best For |
|---|---|---|
mubert/mubert-ttm | Music | AI-generated music and soundtracks |
Deepgram
Speech-to-text.
Capabilities: Speech-to-Text
Featured Models
| Model | Type | Best For |
|---|---|---|
deepgram/nova-3 | STT | General-purpose transcription |
Stability AI
Stable Diffusion image generation models.
Capabilities: Image Generation
Featured Models
| Model | Type | Best For |
|---|---|---|
stability/stable-image-ultra | Image | Highest-quality image generation |
stability/sd3.5-large | Image | High-quality image generation |
stability/sd3.5-large-turbo | Image | Fast image generation |
Black Forest Labs
FLUX models for image generation.
Capabilities: Image Generation
Featured Models
| Model | Type | Best For |
|---|---|---|
bfl/flux-2-max | Image | Highest-quality FLUX.2 images |
bfl/flux-2-pro | Image | Professional quality |
bfl/flux-2-flex | Image | Adjustable quality and speed |
bfl/flux-2-klein-9b | Image | Fast generation |
bfl/flux-kontext-pro | Image | Image editing with context |
Runway
Video generation for creative applications.
Capabilities: Video Generation
Featured Models
| Model | Type | Best For |
|---|---|---|
runway/gen4.5 | Video | Text or image to video |
runway/gen4_turbo | Video | Fast image to video |
runway/aleph2 | Video | Video editing (video to video) |
Luma AI
AI video generation with Ray models.
Capabilities: Video Generation
Featured Models
| Model | Type | Best For |
|---|---|---|
luma/ray-2 | Video | High-quality video |
luma/ray-flash-2 | Video | Fast video generation |
Self-Hosted
Self-hosted models deployed on your own infrastructure. 169 pre-configured models are available across all AI capabilities.
Provider: self_hosted (always -- the engine field determines the API protocol)
Engines:
| Engine | Models | Used For |
|---|---|---|
vllm | 113 | LLM chat, multimodal, embedding, image generation |
transformers | 24 | Custom HuggingFace models not supported by vLLM |
tts-engine | 16 | Text-to-speech models |
whisper | 5 | Speech-to-text models |
video-engine | 6 | Video generation models |
custom | 5 | Generic custom inference servers |
Capabilities by Engine:
| Capability | vllm/transformers | whisper | tts-engine | video-engine | custom |
|---|---|---|---|---|---|
| Chat completions | Yes | - | - | - | Yes |
| Streaming | Yes | - | - | - | Yes |
| Embeddings | Yes | - | - | - | Yes |
| Vision/Multimodal | Yes | - | - | - | - |
| Text-to-Speech | - | - | Yes | - | Yes |
| Speech-to-Text | - | Yes | - | - | Yes |
| Image Generation | Yes | - | - | - | Yes |
| Video Generation | - | - | - | Yes | Yes |
| Tool Calling | Yes | - | - | - | - |
See Self-Hosted Models for the full engine architecture and model catalogue.
Using Models
When calling the AI Gateway, use the Strongly-generated model ID - the unique identifier assigned when a model is added to your account.
// Get your model ID from your configured models list
const modelId = '507f1f77bcf86cd799439011';
const response = await fetch('https://ai-gateway.strongly.ai/v1/chat/completions', {
method: 'POST',
headers: {
'Content-Type': 'application/json',
'X-User-Id': userId,
'X-App-Id': appId // or X-Workflow-Id or X-Workspace-Id
},
body: JSON.stringify({
model: modelId, // Strongly-generated model ID
messages: [{ role: 'user', content: 'Hello!' }]
})
});
Web Search
Add web_search_options to a chat request and the model searches the web before it answers, citing the pages it used. Send {} for the defaults, or {"max_uses": 3} to cap how many searches a request runs.
body: JSON.stringify({
model: modelId,
messages: [{ role: 'user', content: 'What is the latest stable Python release?' }],
web_search_options: {}
})
Which search a model uses depends on the model:
| Models | Search |
|---|---|
| OpenAI: GPT-6 Astra, GPT-6.1 Sol, GPT-6 Sol and GPT-6 Luna, GPT-5 through GPT-5.6 (with their mini, nano and pro models, GPT-5.3 Codex and the GPT-5 search models), GPT-4o, GPT-4o mini, GPT-4.1, GPT-4.1 mini, o3 and o3-pro | OpenAI's own web search |
| Anthropic: Claude Opus 5.5, Sonnet 5.5, Haiku 5.5, Fable 5.1 and every earlier Claude model still served (Haiku 4.5, Sonnet 4.5 and later, Opus 4.5 and later, Fable 5) | Anthropic's own web search |
Google: Gemini 3.8, 3.7, 3.6 and 3.5 Flash, 3.5 Flash-Lite, 3.1 Flash-Lite, 3 Flash, 3.1 Pro, Robotics-ER 2, 2.5 Flash, 2.5 Pro, 2.5 Flash-Lite, and the flash-latest, flash-lite-latest and pro-latest aliases | Google Search grounding |
| Every other model that can call tools: other providers, self-hosted models, and OpenAI models without a search of their own (the Codex and chat-latest models) | Strongly's search, run by the AI Gateway |
A model that can neither search nor call tools (for example an audio-only model) is refused a web search request with the reason; it never answers as if it had searched.
An OpenAI reasoning model asked to search at minimal reasoning effort searches at low, the lowest effort OpenAI's search runs at.
Switching Providers
To switch between providers, use the Strongly ID for the model you want to use:
// Each model has its own unique Strongly ID
const openaiModel = '507f1f77bcf86cd799439011'; // Your GPT-6.1 Sol instance
const anthropicModel = '507f1f77bcf86cd799439012'; // Your Claude instance
const googleModel = '507f1f77bcf86cd799439013'; // Your Gemini instance
// Use any model by passing its Strongly ID
const model = openaiModel;
Note: The same vendor model (e.g., GPT-6.1 Sol) can be configured multiple times with different API keys. Each configuration gets a unique Strongly ID.