ELSEIF
Your brief EB
1,996 stories from 225 feeds 1253 clusters Refreshed 4 minutes ago next pull 00:11

INFRA Signal 98

How to scale Alloy as a central telemetry gateway: capacity planning, load testing, and production lessons

Illustration only Photo by Quilia on Unsplash

elseif has not written about this yet · Grafana describes it this way

Running Alloy as a single-instance sidecar is simple. Running it as a centralized gateway that absorbs the full telemetry stream of an enterprise platform—tens of millions of active series, terabytes of logs per day, and tens of thousands of trace spans per second—is a different challenge altogether. To get it right, you need deliberate capacity planning, honest load testing, and a monitoring setup that doesn't rely on the very thing you're testing.As part of the Professional Services team here at Grafana Labs, we've seen this firsthand working with customers. In this post, we'll walk you through the best practices we follow to help them find success, and we'll do so using real, anonymized data from a recent engagement. We'll cover how we sized and load tested a production Alloy central collector deployment on Kubernetes, what the numbers looked like under real stress, and how the cluster behaves today handling the full production telemetry workload for a large enterprise platform. By the end, you should have a better sense for how you can create your own central gateway for collecting telemetry in Grafana Cloud.Why a central gateway?Before diving into numbers, it's worth explainin
Grafana ↗

THE CLUSTER

Same story, 1 feed.

ORDERED BY FIRST SEEN
Grafana How to scale Alloy as a central telemetry gateway: capacity planning, load testing, and production lessons Open ↗