Skip to main content

AI Analytics

Retrieve usage and performance analytics for AI models managed by the Strongly AI Gateway, counted in requests and tokens. All analytics endpoints are read-only and cover one date range: the last 24h, 7d, 30d or 90d.

Each figure counts the requests the caller may see: a platform administrator sees every request, an organization owner or admin on a multi-tenant platform sees their organization's members' requests, and anyone else sees their own.


GET /api/v1/ai/analytics/usage​

Get usage statistics

Returns the range's totals: requests, tokens, latency, error rate, users and models.

Scope: ai-gateway:read

Parameters:

NameInTypeRequiredDescription
dateRangequerystringYes24h, 7d, 30d or 90d
modelIdquerystringNoOnly this model's requests

Response: 200 OK

{
"data": {
"total_requests": 145230,
"total_tokens": 28500000,
"prompt_tokens": 18200000,
"completion_tokens": 10300000,
"avg_response_time": 485,
"min_response_time": 42,
"max_response_time": 9120,
"p95_response_time": 1500,
"p95_over_bound": false,
"p99_response_time": 3000,
"p99_over_bound": false,
"error_rate": 1.5,
"success_rate": 98.5,
"unique_users_count": 38,
"unique_models_count": 7
},
"meta": { "requestId": "req_abc123" }
}

See UsageStats for the fields.


GET /api/v1/ai/analytics/performance​

Get performance metrics

Returns request latency per model over the range, most requests first.

Scope: ai-gateway:read

Parameters:

NameInTypeRequiredDescription
dateRangequerystringYes24h, 7d, 30d or 90d
modelIdquerystringNoOnly this model

Response: 200 OK

{
"data": [
{
"model_id": "model_abc123",
"model_name": "GPT-4 Production",
"provider": "openai",
"total_requests": 52000,
"avg_response_time": 420,
"min_response_time": 38,
"max_response_time": 8800,
"p95_response_time": 1500,
"p95_over_bound": false,
"p99_response_time": 3000,
"p99_over_bound": false
}
],
"meta": { "requestId": "req_abc123" }
}
FieldTypeDescription
model_idstringModel ID
model_namestringModel name
providerstringModel provider
total_requestsintegerRequests in the range
avg_response_timeinteger | nullMean latency in milliseconds
min_response_timenumber | nullFastest request in milliseconds
max_response_timenumber | nullSlowest request in milliseconds
p95_response_timeinteger | null95th percentile latency in milliseconds, as the upper bound of its latency bucket
p95_over_boundbooleantrue when the 95th percentile is above the highest bucket bound (10000 ms), so the true value is over p95_response_time
p99_response_timeinteger | null99th percentile latency in milliseconds, as the upper bound of its latency bucket
p99_over_boundbooleanAs p95_over_bound, for the 99th percentile

GET /api/v1/ai/analytics/time-series​

Get time-series data

Returns requests, tokens, mean latency and errors per provider, per hour or per UTC day, suitable for charting and trend analysis.

Scope: ai-gateway:read

Parameters:

NameInTypeRequiredDescription
dateRangequerystringYes24h, 7d, 30d or 90d
granularityquerystringYeshourly or daily. hourly is allowed only with 24h; any other range is refused with the message An hourly series is of the last 24 hours only
providerquerystringNoOnly this provider (all means every provider)

Response: 200 OK

{
"data": [
{
"timestamp": "2025-02-06",
"provider": "anthropic",
"requests": 1850,
"tokens": 410000,
"avg_latency": 540,
"errors": 12
},
{
"timestamp": "2025-02-06",
"provider": "openai",
"requests": 4200,
"tokens": 820000,
"avg_latency": 410,
"errors": 30
}
],
"meta": { "requestId": "req_abc123" }
}

data holds one point per period and provider, ordered by timestamp then provider. timestamp is YYYY-MM-DD HH:00 for hourly and YYYY-MM-DD for daily (UTC). With daily, every day of the range is present for each provider: a day with no requests has zero requests, tokens and errors and a null avg_latency (milliseconds).


GET /api/v1/ai/analytics/providers​

Get per-provider statistics

Returns aggregated statistics grouped by AI provider, most requests first, useful for comparing providers side by side.

Scope: ai-gateway:read

Parameters:

NameInTypeRequiredDescription
dateRangequerystringYes24h, 7d, 30d or 90d

Response: 200 OK

{
"data": [
{
"provider": "openai",
"models_count": 5,
"total_requests": 82000,
"total_tokens": 18500000,
"avg_latency": 420,
"unique_users_count": 31
},
{
"provider": "anthropic",
"models_count": 3,
"total_requests": 45000,
"total_tokens": 12000000,
"avg_latency": 550,
"unique_users_count": 22
}
],
"meta": { "requestId": "req_abc123" }
}
FieldTypeDescription
providerstringProvider name
models_countintegerModels of this provider with requests in the range
total_requestsintegerRequests in the range
total_tokensintegerTokens in the range (prompt + completion)
avg_latencyinteger | nullMean latency in milliseconds
unique_users_countintegerUsers with requests to this provider

Response Models Reference​

UsageStats​

FieldTypeDescription
total_requestsintegerTotal number of inference requests
total_tokensintegerTotal tokens consumed (prompt + completion)
prompt_tokensintegerTotal input/prompt tokens
completion_tokensintegerTotal output/completion tokens
avg_response_timeinteger | nullMean latency in milliseconds (null with no requests)
min_response_timenumber | nullFastest request in milliseconds
max_response_timenumber | nullSlowest request in milliseconds
p95_response_timeinteger | null95th percentile latency in milliseconds, as the upper bound of its latency bucket
p95_over_boundbooleantrue when the 95th percentile is above the highest bucket bound (10000 ms)
p99_response_timeinteger | null99th percentile latency in milliseconds, as the upper bound of its latency bucket
p99_over_boundbooleanAs p95_over_bound, for the 99th percentile
error_ratenumber | nullPercentage of requests that failed (0 - 100), null with no requests
success_ratenumber | nullPercentage of requests that succeeded (0 - 100), null with no requests
unique_users_countintegerUsers with requests in the range
unique_models_countintegerModels with requests in the range