Skip to main content

Monitoring Models

Monitor model usage and performance with the AI Gateway analytics: requests, tokens, latency and errors, per model, provider and user. Usage is counted in requests and tokens; no price is put on tokens. Infrastructure cost, including that of self-hosted models, is in FinOps.

AI Gateway Analytics

Who is counted​

Each figure counts the requests you may count: a platform administrator every request, an organization owner or admin on a multi-tenant platform their organization's members, and anyone else their own requests. The same rule applies to the request figures on the AI Gateway overview and on a model's page, including for a model shared with you, such as the platform's embedding model.

What is counted​

Every call the gateway serves is counted: chats and completions, and embeddings, moderation, reranking, speech, transcription, translation, image, video and music calls. For a call other than a chat or a completion, only its figures are kept (model, user, time, response time, tokens, success), never its inputs or outputs. Tokens a call does not report (speech, images, video, music) count as none.

Metrics​

MetricDescription
RequestsCalls to the gateway
TokensInput (prompt) and output (completion) tokens, and their total
Avg Response TimeMean latency, in milliseconds
P95 / P99 Response TimeThe latency 95% / 99% of requests completed within
Error RatePercent of requests that failed
Success RatePercent of requests that succeeded
Users / ModelsDistinct users and models with requests

A figure with no requests behind it (the latency of a range with no requests) is empty, not zero.

Performance Monitoring

Monitor P95 and P99 latency to ensure consistent user experience. High P95/P99 values indicate performance issues affecting some users even if average latency is acceptable.

Analytics API​

The same figures are in the REST API. Every endpoint takes a required dateRange: 24h, 7d, 30d or 90d, the range ending now.

EndpointReturnsOther parameters
GET /api/v1/ai-gateway/analytics/usageThe range's totals: requests, tokens, latency, error and success rates, users, modelsmodelId
GET /api/v1/ai-gateway/analytics/performancePer model: requests and latency (average, minimum, maximum, P95, P99)modelId
GET /api/v1/ai-gateway/analytics/time-seriesPer provider and period: requests, tokens, mean latency, errorsgranularity (required): hourly (with 24h only) or daily; provider
GET /api/v1/ai-gateway/analytics/providersPer provider: models, requests, tokens, mean latency, users

A daily series has every day of the range, with zero requests on a day without any. See the AI Analytics API and the Python SDK's client.ai.analytics.

Guardrails Monitoring​

Platform-Wide Guardrail Metrics​

  • Total Models: Number of deployed models (self-hosted + third-party)
  • Protected Models: Count of models with guardrails enabled
  • Total Rules: Aggregate count of all active guardrail rules
  • Blocked Today: Requests blocked by guardrails in last 24 hours
  • Modified Today: Requests modified by guardrails in last 24 hours

Per-Model Guardrail Metrics​

Each model tracks:

  • Enabled Status: Whether guardrails are active for this model
  • Rules Count: Total number of configured rules
  • Input Rules: Rules applied to user prompts before model inference
  • Output Rules: Rules applied to model responses before returning to user
  • Total Requests: Number of API calls processed
  • Blocked Requests: Requests denied due to policy violations
  • Modified Requests: Requests altered by guardrails (e.g., PII redaction)
  • Block Rate: Percentage of requests blocked

Guardrail Rule Types​

Monitor specific rule categories:

Content Filtering​

  • Toxicity Detection: Blocks toxic, offensive, or harmful content

    • Categories: Hate speech, harassment, violence, profanity, sexual content
    • Threshold levels: Low (permissive), Medium (balanced), High (strict)
  • PII Detection: Identifies and redacts personally identifiable information

    • Detects: Email addresses, phone numbers, SSN, credit cards, IP addresses, names, addresses, dates of birth
    • Actions: Redact, mask, or block entire request
  • Prompt Injection Detection: Detects attempts to manipulate model behavior

    • Patterns: "Ignore previous instructions", system prompt extraction, role confusion

Topic Restrictions​

  • Allowed Topics: Restricts model to specific subject areas
  • Banned Topics: Blocks specific prohibited subjects

Output Validation​

  • Format Enforcement: Ensures output matches required structure
    • Formats: JSON schema, XML, markdown, specific patterns
    • Action: Retry generation or return error

Filters & Options​

Customize your analytics view:

Date Range​

  • Last 24 hours
  • Last 7 days
  • Last 30 days
  • Last 90 days

Provider Filter​

View all providers or filter by:

  • OpenAI
  • Anthropic
  • Google (Gemini)
  • Mistral
  • Cohere
  • DeepSeek
  • Grok (xAI)
  • ElevenLabs
  • Stability AI
  • Black Forest Labs
  • Runway
  • Luma AI
  • Self-Hosted (vLLM)

API Access​

All analytics data is available programmatically via the REST API endpoints listed above.

Best Practices​

Regular Monitoring​

  • Daily: Check summary metrics and traffic trends
  • Weekly: Review model performance and optimize configurations
  • Monthly: Analyze usage patterns and forecast future needs

Performance Optimization​

  1. Identify Slow Models: Sort by P95/P99 latency
  2. Analyze Error Rates: Investigate models with high error rates
  3. Optimize Token Usage: Review input/output token ratios
  4. Scale Appropriately: Adjust autoscaling based on usage patterns

Infrastructure Cost​

Self-hosted models cost what their compute costs; see FinOps for that cost and for budgets.

Security Monitoring​

  1. Review Guardrail Activity: Check blocked and modified requests
  2. Investigate Anomalies: Look for unusual traffic patterns
  3. Monitor User Activity: Track per-user usage for abuse detection
  4. Update Rules: Adjust guardrails based on observed patterns

Troubleshooting​

High Error Rates​

Possible causes:

  • Model overloaded (needs more resources)
  • API key issues with third-party provider
  • Network connectivity problems
  • Invalid request formats

Solutions:

  • Enable autoscaling or add more instances
  • Verify API key validity
  • Check network connectivity
  • Review request logs for formatting issues

High Latency​

Possible causes:

  • Insufficient compute resources
  • Cold starts (on-demand deployment)
  • Large input/output token counts
  • Network latency

Solutions:

  • Scale up resources or enable autoscaling
  • Use "Always On" deployment
  • Optimize prompts to reduce token usage
  • Deploy in regions closer to users

Unexpected Infrastructure Cost​

Possible causes:

  • Self-hosted models autoscaling to max replicas
  • Forgotten "Always On" self-hosted models

Solutions:

  • Review autoscaling configuration
  • Audit deployed models and stop unused ones
  • Check the models' cost in FinOps

Guardrails Over-Blocking​

Possible causes:

  • Thresholds too strict
  • False positives in content detection
  • Overly broad topic restrictions

Solutions:

  • Adjust sensitivity thresholds
  • Review blocked requests to identify patterns
  • Refine allowed/banned topic lists
  • Add exemptions for legitimate use cases

Next Steps​