INFRA Signal 93
Atlassian automates root cause analysis by correlating anomalies across metrics, logs, traces, and service topology
Atlassian detailed a system that automates root cause analysis for cloud-native incidents by correlating anomalies across metrics, logs, traces, and service topology to generate ranked hypotheses.
Manual root cause analysis requires engineers to visually correlate telemetry across separate dashboards, a process that depends heavily on domain knowledge and delays resolution. Automating the hypothesis generation step allows responders to skip straight to validation, reducing the cognitive load during high-pressure incidents. The modular architecture also allows teams to swap anomaly detection models or add new signal types without rebuilding the pipeline.
Written by elseif from the cluster below · every claim links back to a sourceThe three things worth knowing
The system scopes the incident blast radius by querying an OpenTelemetry-derived service dependency graph to identify the affected service subgraph.
Independent anomaly detectors process metrics using median absolute deviation, traces using structural pattern analysis, and logs using embedding-based clustering.
All detectors emit normalized anomaly events that a correlation engine aligns on a shared timeline and traces through the service dependency graph to rank hypotheses.
THE CLUSTER
↗