Cost Optimization
Optimize your AI infrastructure costs with smart deployment strategies, resource configuration, and usage patterns. This guide covers best practices for reducing costs while maintaining performance and reliability.
Understanding AI Model Costs
Self-Hosted Models
Self-hosted models incur compute and storage costs:
- Compute Costs: Pay for GPU/CPU resources while instances are running
- Storage Costs: Pay for model storage and volumes
- Network Costs: Data transfer and load balancing
- Deployment Model: Always On, On Demand (auto-shutdown), or Scheduled
Third-Party Providers
Third-party providers charge per usage:
- Token-Based Pricing: Pay per input and output token
- Model Tier Pricing: Different rates for different models
- Volume Discounts: Some providers offer bulk pricing
- No Infrastructure Costs: Provider manages infrastructure
Cost Optimization Strategies
1. Choose the Right Deployment Model
On Demand with Auto-Shutdown (70-90% Savings)
Best for:
- Development and testing
- Low-traffic applications
- Intermittent workloads
- Prototype and demo applications
How it works:
- Model runs when needed
- Automatically shuts down after configurable idle period (
auto_shutdown_minutes) - Pay only for active time
Configuration:
Deployment: On Demand
Auto Shutdown: 5 minutes idle
Effect: the GPU runs only while the model serves, so a model used a few hours a day costs a fraction of one left running. What a deployment costs is the infrastructure it runs on, shown in FinOps.
Scheduled Deployment
Best for:
- Business hours only (9am-5pm)
- Batch processing jobs
- Regional availability windows
- Predictable usage patterns
How it works:
- Runs only during specified time windows
- Use cron expressions to define schedule
- Automatically starts and stops
Configuration:
Deployment: Scheduled
Schedule: "0 9-17 * * 1-5" # Weekdays 9am-5pm
# or
Schedule: "0 8 * * *" # Daily at 8am
Effect: a weekday business-hours window runs 40 of the week's 168 hours. What a deployment costs is the infrastructure it runs on, shown in FinOps.
Always On
Best for:
- Production applications with consistent traffic
- Low-latency requirements
- High availability needs
Trade-offs:
- No cold starts
- Instant response
- Fixed costs regardless of usage
2. Right-Size Your Resources
Match GPU to Model Size
| Model Size | Recommended GPU | Oversized GPU |
|---|---|---|
| 7B params | 1x V100 (16GB) | 1x A100 (40GB) |
| 13B params | 1x A10G (24GB) | 1x A100 (40GB) |
| 70B params | 4x A100 (40GB) | 8x A100 (80GB) |
Best practice: Use the smallest GPU that fits your model
Use Quantization
Reduce model size and GPU requirements with quantization:
| Model | Full Precision | 8-bit Quantized | 4-bit Quantized |
|---|---|---|---|
| 70B Model | 4x A100 (160GB) | 2x A100 (80GB) | 1x A100 (40GB) |
| Quality Loss | 0% | Less than 2% | 2-5% |
Effect: half to a quarter of the GPUs, with little quality loss.
3. Optimize Autoscaling Configuration
Conservative Scaling (Avoid Over-Scaling)
Problem: Aggressive autoscaling can cause unnecessary scaling:
- Scaling up on temporary spikes
- Not scaling down fast enough
Solution:
CPU Threshold: 75% # Higher threshold
Memory Threshold: 85% # Higher threshold
Scale Down Cooldown: 300s # Longer cooldown to avoid premature scale-up
Right-Size Min/Max Replicas
Problem: Min replicas too high = wasted capacity
Solution:
- Development: Min=0 or 1, Max=3
- Production (moderate): Min=2, Max=10
- Production (high-traffic): Min=3, Max=20
Effect: every minimum replica runs all the time, so 2 instead of 5 runs 60% fewer replicas at idle.
4. Optimize Token Usage
Reduce Prompt Length
Problem: Long prompts use more tokens
Solution:
- Remove unnecessary context
- Use concise instructions
- Implement prompt templates
- Cache common context
Example:
- Verbose prompt: 500 tokens
- Optimized prompt: 150 tokens
- Savings: 70% on input tokens
Limit Output Tokens
Problem: Unlimited outputs can generate excessive tokens
Solution:
- Set
max_tokensparameter - Request concise responses
- Use streaming to stop generation early
Example:
{
model: '507f1f77bcf86cd799439011', // Strongly-generated model ID
messages: [{ role: 'user', content: 'Summarize this article...' }],
max_tokens: 500, // Limit response length
temperature: 0.7
}
Implement Response Caching
Problem: Repeated identical requests waste tokens
Solution:
- Cache frequent queries
- Use Redis or similar
- Set appropriate TTL
Savings: 50-80% for repeated queries
5. Choose Cost-Effective Models
Model Tier Selection
| Task Complexity | Recommended Model | Overkill Model |
|---|---|---|
| Simple classification | GPT-6 Luna | GPT-6 Astra |
| General Q&A | Claude Haiku 5.5 | Claude Opus 5.5 |
| Basic completion | Mistral-7B (self-hosted) | Llama-70B |
Best practice: Use smallest model that meets quality requirements
Self-Hosted vs Third-Party
The platform puts no price on tokens: a third-party model is billed by its provider, per token, on your own provider account, and the AI Gateway counts its usage in requests and tokens. A self-hosted model costs the infrastructure it runs on, shown in FinOps. Compare the provider's bill for your token volume with that infrastructure cost; at low and medium volumes third-party models usually cost less, and self-hosting pays off only with sustained high usage and a right-sized, optimized deployment.
6. Implement Smart Routing
Route by Task Complexity
Strategy: Use different models for different task types
// Use Strongly-generated model IDs from your configured models
const MODELS = {
simple: '507f1f77bcf86cd799439011', // GPT-6 Luna
medium: '507f1f77bcf86cd799439012', // Claude Haiku 5.5
complex: '507f1f77bcf86cd799439013', // GPT-6 Astra
};
function selectModel(task) {
if (task.complexity === 'simple') return MODELS.simple;
if (task.complexity === 'medium') return MODELS.medium;
return MODELS.complex;
}
Savings: 50-80% by avoiding overkill models
7. Monitor and Optimize Continuously
Track Usage
The AI Gateway analytics count usage in requests and tokens (they put no price on tokens). Use them to see which models, providers and users drive usage:
- Requests and tokens per model (
GET /api/v1/ai-gateway/analytics/performanceandusage) - Requests and tokens per provider (
GET /api/v1/ai-gateway/analytics/providers) - Usage over time (
GET /api/v1/ai-gateway/analytics/time-series)
See Monitoring. The infrastructure cost of self-hosted models is in FinOps.
Regular Cost Reviews
Weekly:
- Review top 5 most expensive models
- Identify optimization opportunities
- Check for unused or idle models
Monthly:
- Analyze usage trends
- Adjust autoscaling configurations
- Evaluate model selection
8. Batch Processing
Batch Similar Requests
Problem: Individual requests have overhead
Solution:
- Group similar requests
- Process in batches
- Share context across requests
Savings: 20-40% through reduced overhead
Off-Peak Processing
Problem: On-demand processing during peak hours
Solution:
- Queue non-urgent requests
- Process during off-peak hours
- Use scheduled deployment
Savings: Leverage scheduled deployment savings
Cost Optimization Checklist
Development Phase
- Use on-demand deployment with auto-shutdown
- Start with small models (7B parameters)
- Use third-party providers for prototyping
- Set max_tokens limits
- Disable autoscaling (use 1 instance)
Testing Phase
- Implement response caching
- Test with quantized models
- Optimize prompt templates
- Set up usage monitoring (requests and tokens) via the analytics API
- Use scheduled deployment for batch tests
Production Phase
- Right-size GPU resources
- Configure autoscaling appropriately
- Implement smart routing
- Monitor usage trends via analytics endpoints
- Use model tier appropriate for tasks
- Monitor tokens per request
- Review and optimize monthly
Common Cost Pitfalls
1. Forgotten Always-On Models
Problem: Models left running 24/7 unnecessarily
Solution: Audit all deployments monthly, switch to on-demand if appropriate
Cost impact: a GPU running 24/7 for nothing; find it in FinOps
2. Oversized GPU Selection
Problem: Using A100 for 7B model that fits on V100
Solution: Match GPU to model requirements
Cost impact: the difference between the two GPUs, every hour it runs
3. No Token Limits
Problem: Unlimited output generation
Solution: Set max_tokens based on use case
Cost impact: 2-10x more tokens
4. Autoscaling to Max Unnecessarily
Problem: Aggressive autoscaling hitting max replicas
Solution: Tune thresholds and cooldown periods
Cost impact: Thousands per month in unnecessary scaling
5. Wrong Provider for Volume
Problem: Using self-hosted for low volume
Solution: Use third-party until reaching cost tipping point
Cost impact: 5-10x higher costs at low volume
Next Steps
- Monitor your costs with analytics endpoints
- Configure autoscaling for optimal resource usage
- Review deployment options for your use case