Dashboard
The Dashboard (/dashboard) shows how your workloads are doing: their health, how much they are used, what is failing and what changed. What workloads cost is shown in FinOps only. Open it from Dashboard in the sidebar.
What you see
What the dashboard covers depends on your role:
| Role | What the dashboard covers |
|---|---|
| admin | Every workload on the platform, titled Platform dashboard. An All organizations picker narrows every section to one organization. It lists the organizations that have workloads, by name; two with the same name are told apart by their owner's email. Capacity by node pool appears only for admins. |
| developer | Only the workloads you own, titled Your dashboard. Figures about people are about the users of your apps. |
| app | No dashboard. Opening it takes you to My Apps. |
The platform enforces this when it reads the data, not only on the page: a developer's request never returns another user's workloads or incidents.
A workload is anything you deploy and run. The dashboard groups workloads by type:
| Type | Primary use | Opens |
|---|---|---|
| Apps | Distinct users | /apps |
| Workflows | Runs | /workflow-builder |
| Agents | Runs | /agents |
| Jobs | Runs (scheduled and manual project jobs) | /projects |
| GenAI models | Requests through the AI Gateway | /ai-gateway |
| Workspaces | Active hours (time someone was using the IDE) | /workspaces |
| Model registry | Predictions | /mlops/model-registry |
| Add-ons | Operations reported by the database or service engine | /addons |
Time window
The 24h, 7d and 30d buttons choose the window. Use (users, runs, requests, active hours, predictions, operations), incidents and changes cover the last 24 hours, 7 days or 30 days up to now. Each change badge compares a figure with the window of the same length just before it. Trends are hourly for 24 hours and daily for 7 and 30 days; the first and last points can cover part of an hour or day.
App users appear within about two minutes of using an app. A user counts in every window and trend point their visit overlaps, however long ago the visit began. Every way of opening an app counts: View App, the app's link, branded sign-in and Home App, including its streaming and websocket traffic. A visitor of a public link who is not signed in can't be identified, so they count in the app's requests but not as a user.
A figure the platform has not measured is shown as not measured, never as 0. For example, an add-on whose engine does not report operations shows not measured for its use.
Health
Every workload has one of four health states. The dashboard, the incident list and the per-type pages all use the same definition:
| Health | Meaning |
|---|---|
| Failing | The workload is in error (for example, a failed deployment), or one of its pods cannot start. |
| Degraded | It is running, but a pod is not ready or has restarted. A workflow, agent or job is also degraded when the share of its runs that failed recently is above the platform's threshold. |
| Healthy | It is running with none of the above. |
| Stopped | It is deployed but not running: stopped, scaled to zero, idle, or never deployed. |
A job is Running while one of its runs is pending or running, and Failing when its latest run failed or one of its pods cannot start. Between runs it is Stopped, unless too many of its recent runs failed (Degraded).
The thresholds are platform settings: the failed-run share (20% by default) over a recent window (24 hours by default). Pod readiness comes from the platform's pod readings, and Fleet at a glance shows the time of the oldest reading it used.
Sections
Your budgets
Developers and users see Your budgets above the other sections when a FinOps budget covers them: their own user budget, their organization's, or the platform's. Each budget shows what remains, the days until it resets, and a warning when it is exhausted (new launches are then blocked). The card appears only when FinOps is enabled and at least one budget applies.
Headline figures
Four figures across the top:
- Running workloads: healthy plus degraded workloads, out of the total, with the number failing and stopped.
- Active users (admins): users active today, this week and this month, and stickiness (daily as a share of monthly). Users of your apps (developers): the distinct users of your apps in the window, the change, and the share who came back.
- Open incidents (admins) or Open issues (developers): how many are open, by severity, and how many workloads they hit.
- Capacity in use (admins): what the pods on Ready nodes request, as a share of what those nodes can allocate (CPU, then memory and GPUs). Pods that no node holds yet are demand, not use: they are counted apart as waiting for capacity, with what they request. Your requested capacity (developers): the CPU, memory and GPUs your running pods request, the share of that CPU actually used, and any of your pods waiting for capacity.
Fleet at a glance
One row per workload type, showing:
- workloads Running, Failing and Degraded;
- Primary use, with its change and trend;
- open Incidents.
Each column is shaded on its own scale, so the outliers stand out. Select a row to open that type's page (see Workload type pages).
Needs attention and What changed
Needs attention lists open incidents, widest first. An incident is a group of failures with the same cause, such as one error message across several workloads. Each incident shows:
- its severity and source (Cannot start, Failed or Failing runs) and when it began;
- an example message;
- the workloads it hits, how many owners they belong to (admins), and how often it occurred;
- what those workloads depend on.
A Cannot start or Failed incident stays open while any of its workloads still fails. A Failing runs incident is resolved when no run has failed for a quiet period (60 minutes by default). Incidents resolved within the window stay listed, marked Resolved.
What changed lists deploys, updates, configuration changes, starts, stops, deletions and platform upgrades, newest first. A change is linked to an incident when it touched one of the incident's workloads or their dependencies within the quiet period before the incident began. The incident then shows how many changes came before it, and pointing at the incident highlights those changes.
Most used and Biggest movers
Most used has a tab per workload type. Each tab ranks the workloads by their primary use, against the window before, with:
- Apps: sessions and average session length;
- Workflows: failed runs and average duration;
- GenAI models: calling users and failed requests.
Biggest movers shows the largest rises in use between the two windows, or the largest falls (choose Rising or Falling; each shows how many moved), 10 at a time with Previous and Next, and how many workloads were newly used.
Top creators and Adoption
For admins, Top creators ranks owners by the distinct users of the apps they built, then by their workflows' runs. Adoption shows active users over time and the workloads created in the window.
For developers, Who uses your workloads lists the users of your apps by sessions. Adoption shows the users of your apps and the workloads you created.
Idle workspaces and capacity
- Idle workspaces: running workspaces nobody has used for the idle period (4 hours by default), never used first. Each shows when it was last used; a workspace with no recorded use shows no activity recorded since the date the platform began recording workspace use.
- Capacity by node pool (admins): per node pool, the number of nodes (spot nodes noted), and the CPU, memory and GPUs requested by the pods on its Ready nodes out of what those nodes can allocate. A bar turns amber at 75% and red at 90%. Pods waiting for capacity are listed apart, with what they request.
Workload type pages
Selecting a row in Fleet at a glance opens that type's page at /dashboard/<type> (for example /dashboard/jobs). Admins see every workload of the type and can narrow the page to one organization; developers see their own, marked Yours. The window buttons and the type buttons under the title switch the window and the type without leaving the page. Open the ... page goes to the type's own management page.
Each page has, from the top:
-
KPIs: the type's figures, each against the previous window where it has one, with a trend. For example, Jobs shows how many are scheduled, runs, success rate, average duration and failing jobs. A figure that needs runs or calls says so when there were none (for example none in the window or no tool calls) rather than not measured.
-
Health over time: how many workloads were degraded and failing in each hour or day, and the top failure reasons in the window. Health is recorded every 5 minutes, so the chart starts when recording began.
-
Incidents: the type's incidents grouped by cause, with how many are open, critical and resolved, and the median time to resolve.
-
Usage: the type's use over the window and in numbers (for Jobs: runs succeeded and failed, success rate, cancelled, in progress, average duration).
-
The type's own panels:
Type Panels Apps When apps are used: sessions started by weekday and hour (UTC) Workflows Failing nodes by node type; runs by trigger; the dead-letter queue by workflow Agents Tool errors by tool; run time (p50 and p95); tokens by agent (counted, never priced) Jobs Next runs, soonest first; failed runs in the window, newest first, with why each failed GenAI models Failed requests by cause; guardrail hits by rule and what the rule did; top callers Workspaces Running workspaces by IDE Model registry Drift over the last 7 days (each model's worst status per day); A/B experiments; prediction latency by model Add-ons Storage: how full each add-on's volumes are, fullest first, from each node's latest reading (a node that could not be read says why); add-ons by engine, with how many run; data sources by type, with those failing their connection test -
Most used: the workloads ranked by use, against the window before.
-
Top creators (admins): owners, or organizations, ranked by the use of the workloads they own, with how ownership is spread. Who uses your ... (developers): the users of your workloads.
-
All ...: every workload of the type, those that need attention (failing or degraded) first. Filter by health (each filter shows its count), search by name, owner or organization, and sort by severity, use, last change or name. Admins can also filter by owner. A failing or degraded workload says why, and each shows its last change. The list is paged, 10 workloads per page.
On Model registry, a model that is idle (scaled to zero; its next request wakes it) counts as deployed.
Customize the dashboard
Select Customize to choose how the dashboard is laid out:
- Presets:
- Overview shows everything, for the last 24 hours.
- Ops morning check puts incidents and changes first and hides people.
- Sections: show, hide or reorder the six sections.
- Workload types in the fleet: choose which types Fleet at a glance shows.
- Views: save the current layout and window under a name, and reopen or delete it later. Reset returns to the standard layout.
Your layout and saved views are kept in your browser, for your account. They do not follow you to another browser or device.
Select the refresh button to reload the figures.