INFRA Signal 440
Reported incident counts rise as teams improve detection and response processes
An increase in reported incidents may reflect better incident management practices rather than declining system reliability.
Engineering teams often use incident counts as a proxy for system reliability, but this metric can mislead. A rising number of reported incidents may signal improved detection and response capabilities, not worsening system health. Misinterpreting this trend could discourage transparency and hinder operational improvements.
Written by elseif from the cluster below · every claim links back to a sourceThe three things worth knowing
Higher incident counts often result from better tooling, training, and willingness to declare issues formally.
Incident volume measures an organization’s ability to surface problems, not underlying system reliability.
Discouraging incident declarations to reduce counts can hide problems and delay critical responses.
THE READ
What the cluster adds up to.
The assumption that more incidents mean less reliability is widespread but flawed. When teams invest in better incident management, such as improved tooling, structured postmortems, and clearer escalation paths, they become more likely to formally declare issues that were previously handled informally. This shift increases reported incident counts, but it also provides greater visibility into operational challenges. The change reflects a cultural and procedural improvement, not a decline in system stability.
Incident counts alone are a poor indicator of reliability because they measure activity, not impact. A team that declares more incidents may actually be reducing risk by addressing problems earlier and learning from each event. Conversely, a team that suppresses incident reports to keep counts low may be masking systemic issues, leading to larger outages later. The metric’s limitations become clear when compared to other engineering disciplines: better vulnerability scanning or observability tools often reveal more issues, not because systems are failing more, but because detection has improved.
Modern reliability engineering emphasizes metrics that reflect customer impact and organizational learning over raw incident counts. Service Level Objectives (SLOs), error budget burn, and user-centric measures like SLI degradation provide a clearer picture of system health. These approaches focus on how failures affect users and how quickly teams recover, rather than the sheer volume of incidents. Shifting focus to these metrics can help teams avoid the pitfalls of over-indexing on incident counts, which can create perverse incentives to hide problems.
The cultural implications of this shift are significant. If engineers believe they will be penalized for declaring incidents, they may delay escalations or attempt to resolve issues alone, reducing organizational visibility. Encouraging incident declarations, even when counts rise, fosters transparency and collective problem-solving. This aligns with broader trends in reliability engineering, where the goal is not to eliminate incidents but to build systems and processes that handle them effectively and learn from them.
Written by elseif from the cluster below · checked for specifics the sources never containedTHE CLUSTER
↗