Reliability
Reliability covers the two surfaces that tell you when something needs your attention: Errors, the technical failures the agents hit, and Incidents, the higher-severity security and governance matters that carry their own clock.

Errors
Errors groups recurring failures by their signature rather than showing you one failure at a time, sorted by how often each one happens over the last 60 days, with a daily trend chart above the table. A glance band totals your errors, how many distinct signatures you're seeing, how many were auto-remediated, with a percentage of the total, and how many were escalated.
Work this list by signature: fix the underlying cause once and the whole group clears, rather than chasing individual failures one at a time.
Incidents
Incidents is the higher bar: it holds security and governance matters that an alert opened automatically, or that an operator declared by hand. Each one carries a severity and a live timer, running against a fixed 15-minute containment target and a 48-hour customer-notification target. A glance band shows Open incidents, Open P1s, how many have breached their target, and how many closed recently. Two lists follow: open incidents with their timers running, and recently closed incidents with the root cause recorded against each.
Work the open list by severity, starting with anything close to breaching its target. Closing an incident means recording what caused it; that stops the clock and keeps the resolution on the record for later review.
Related
- System Status: agent activity and health, for the everyday view.
- Platform Event Log: the wider event stream errors and incidents are drawn from.