Skip to main content

Cost Optimization

Optimize your AI infrastructure costs with smart deployment strategies, resource configuration, and usage patterns. This guide covers best practices for reducing costs while maintaining performance and reliability.

Understanding AI Model Costs​

Self-Hosted Models​

Self-hosted models incur compute and storage costs:

  • Compute Costs: Pay for GPU/CPU resources while instances are running
  • Storage Costs: Pay for model storage and volumes
  • Network Costs: Data transfer and load balancing
  • Deployment Model: Always On, On Demand (auto-shutdown), or Scheduled

Third-Party Providers​

Third-party providers charge per usage:

  • Token-Based Pricing: Pay per input and output token
  • Model Tier Pricing: Different rates for different models
  • Volume Discounts: Some providers offer bulk pricing
  • No Infrastructure Costs: Provider manages infrastructure

Cost Optimization Strategies​

1. Choose the Right Deployment Model​

On Demand with Auto-Shutdown (70-90% Savings)​

Best for:

  • Development and testing
  • Low-traffic applications
  • Intermittent workloads
  • Prototype and demo applications

How it works:

  • Model runs when needed
  • Automatically shuts down after configurable idle period (auto_shutdown_minutes)
  • Pay only for active time

Configuration:

Deployment: On Demand
Auto Shutdown: 5 minutes idle

Effect: the GPU runs only while the model serves, so a model used a few hours a day costs a fraction of one left running. What a deployment costs is the infrastructure it runs on, shown in FinOps.

Scheduled Deployment​

Best for:

  • Business hours only (9am-5pm)
  • Batch processing jobs
  • Regional availability windows
  • Predictable usage patterns

How it works:

  • Runs only during specified time windows
  • Use cron expressions to define schedule
  • Automatically starts and stops

Configuration:

Deployment: Scheduled
Schedule: "0 9-17 * * 1-5" # Weekdays 9am-5pm
# or
Schedule: "0 8 * * *" # Daily at 8am

Effect: a weekday business-hours window runs 40 of the week's 168 hours. What a deployment costs is the infrastructure it runs on, shown in FinOps.

Always On​

Best for:

  • Production applications with consistent traffic
  • Low-latency requirements
  • High availability needs

Trade-offs:

  • No cold starts
  • Instant response
  • Fixed costs regardless of usage

2. Right-Size Your Resources​

Match GPU to Model Size​

Model SizeRecommended GPUOversized GPU
7B params1x V100 (16GB)1x A100 (40GB)
13B params1x A10G (24GB)1x A100 (40GB)
70B params4x A100 (40GB)8x A100 (80GB)

Best practice: Use the smallest GPU that fits your model

Use Quantization​

Reduce model size and GPU requirements with quantization:

ModelFull Precision8-bit Quantized4-bit Quantized
70B Model4x A100 (160GB)2x A100 (80GB)1x A100 (40GB)
Quality Loss0%Less than 2%2-5%

Effect: half to a quarter of the GPUs, with little quality loss.

3. Optimize Autoscaling Configuration​

Conservative Scaling (Avoid Over-Scaling)​

Problem: Aggressive autoscaling can cause unnecessary scaling:

  • Scaling up on temporary spikes
  • Not scaling down fast enough

Solution:

CPU Threshold: 75%  # Higher threshold
Memory Threshold: 85% # Higher threshold
Scale Down Cooldown: 300s # Longer cooldown to avoid premature scale-up

Right-Size Min/Max Replicas​

Problem: Min replicas too high = wasted capacity

Solution:

  • Development: Min=0 or 1, Max=3
  • Production (moderate): Min=2, Max=10
  • Production (high-traffic): Min=3, Max=20

Effect: every minimum replica runs all the time, so 2 instead of 5 runs 60% fewer replicas at idle.

4. Optimize Token Usage​

Reduce Prompt Length​

Problem: Long prompts use more tokens

Solution:

  • Remove unnecessary context
  • Use concise instructions
  • Implement prompt templates
  • Cache common context

Example:

  • Verbose prompt: 500 tokens
  • Optimized prompt: 150 tokens
  • Savings: 70% on input tokens

Limit Output Tokens​

Problem: Unlimited outputs can generate excessive tokens

Solution:

  • Set max_tokens parameter
  • Request concise responses
  • Use streaming to stop generation early

Example:

{
model: '507f1f77bcf86cd799439011', // Strongly-generated model ID
messages: [{ role: 'user', content: 'Summarize this article...' }],
max_tokens: 500, // Limit response length
temperature: 0.7
}

Implement Response Caching​

Problem: Repeated identical requests waste tokens

Solution:

  • Cache frequent queries
  • Use Redis or similar
  • Set appropriate TTL

Savings: 50-80% for repeated queries

5. Choose Cost-Effective Models​

Model Tier Selection​

Task ComplexityRecommended ModelOverkill Model
Simple classificationGPT-6 LunaGPT-6 Astra
General Q&AClaude Haiku 5.5Claude Opus 5.5
Basic completionMistral-7B (self-hosted)Llama-70B

Best practice: Use smallest model that meets quality requirements

Self-Hosted vs Third-Party​

The platform puts no price on tokens: a third-party model is billed by its provider, per token, on your own provider account, and the AI Gateway counts its usage in requests and tokens. A self-hosted model costs the infrastructure it runs on, shown in FinOps. Compare the provider's bill for your token volume with that infrastructure cost; at low and medium volumes third-party models usually cost less, and self-hosting pays off only with sustained high usage and a right-sized, optimized deployment.

6. Implement Smart Routing​

Route by Task Complexity​

Strategy: Use different models for different task types

// Use Strongly-generated model IDs from your configured models
const MODELS = {
simple: '507f1f77bcf86cd799439011', // GPT-6 Luna
medium: '507f1f77bcf86cd799439012', // Claude Haiku 5.5
complex: '507f1f77bcf86cd799439013', // GPT-6 Astra
};

function selectModel(task) {
if (task.complexity === 'simple') return MODELS.simple;
if (task.complexity === 'medium') return MODELS.medium;
return MODELS.complex;
}

Savings: 50-80% by avoiding overkill models

7. Monitor and Optimize Continuously​

Track Usage​

The AI Gateway analytics count usage in requests and tokens (they put no price on tokens). Use them to see which models, providers and users drive usage:

  • Requests and tokens per model (GET /api/v1/ai-gateway/analytics/performance and usage)
  • Requests and tokens per provider (GET /api/v1/ai-gateway/analytics/providers)
  • Usage over time (GET /api/v1/ai-gateway/analytics/time-series)

See Monitoring. The infrastructure cost of self-hosted models is in FinOps.

Regular Cost Reviews​

Weekly:

  • Review top 5 most expensive models
  • Identify optimization opportunities
  • Check for unused or idle models

Monthly:

  • Analyze usage trends
  • Adjust autoscaling configurations
  • Evaluate model selection

8. Batch Processing​

Batch Similar Requests​

Problem: Individual requests have overhead

Solution:

  • Group similar requests
  • Process in batches
  • Share context across requests

Savings: 20-40% through reduced overhead

Off-Peak Processing​

Problem: On-demand processing during peak hours

Solution:

  • Queue non-urgent requests
  • Process during off-peak hours
  • Use scheduled deployment

Savings: Leverage scheduled deployment savings

Cost Optimization Checklist​

Development Phase​

  • Use on-demand deployment with auto-shutdown
  • Start with small models (7B parameters)
  • Use third-party providers for prototyping
  • Set max_tokens limits
  • Disable autoscaling (use 1 instance)

Testing Phase​

  • Implement response caching
  • Test with quantized models
  • Optimize prompt templates
  • Set up usage monitoring (requests and tokens) via the analytics API
  • Use scheduled deployment for batch tests

Production Phase​

  • Right-size GPU resources
  • Configure autoscaling appropriately
  • Implement smart routing
  • Monitor usage trends via analytics endpoints
  • Use model tier appropriate for tasks
  • Monitor tokens per request
  • Review and optimize monthly

Common Cost Pitfalls​

1. Forgotten Always-On Models​

Problem: Models left running 24/7 unnecessarily

Solution: Audit all deployments monthly, switch to on-demand if appropriate

Cost impact: a GPU running 24/7 for nothing; find it in FinOps

2. Oversized GPU Selection​

Problem: Using A100 for 7B model that fits on V100

Solution: Match GPU to model requirements

Cost impact: the difference between the two GPUs, every hour it runs

3. No Token Limits​

Problem: Unlimited output generation

Solution: Set max_tokens based on use case

Cost impact: 2-10x more tokens

4. Autoscaling to Max Unnecessarily​

Problem: Aggressive autoscaling hitting max replicas

Solution: Tune thresholds and cooldown periods

Cost impact: Thousands per month in unnecessary scaling

5. Wrong Provider for Volume​

Problem: Using self-hosted for low volume

Solution: Use third-party until reaching cost tipping point

Cost impact: 5-10x higher costs at low volume

Next Steps​