All pipelines operational
All insights
Platform EngineeringObservability guide

Kubernetes Monitoring: Configuring Prometheus and Grafana for Scale

A single Prometheus scraping a small cluster needs almost no thought. Thirty clusters and several thousand workloads later, the same configuration falls over — not because Prometheus is slow, but because label cardinality grows multiplicatively and nobody set a budget. Scaling Kubernetes monitoring is mostly the discipline of deciding what not to store.

Treat cardinality as a budget

Every unique label combination is a distinct time series held in memory. Pod names, request IDs, user identifiers, and full URL paths are the usual offenders. Drop or aggregate them with relabeling rules at scrape time, before they ever reach storage.

Monitor your monitoring: track series count per job and alert when a team's footprint grows abnormally. A single bad label added in a library upgrade can double a cluster's series count overnight.

Choose a topology that matches the fleet

Run one Prometheus per cluster for local scraping and alerting, then ship aggregated data to a long-term store such as Thanos, Mimir, or Cortex via remote write. Local instances keep short retention and stay fast; the central store handles global queries and long retention.

Federation is acceptable for pulling a small set of aggregated metrics upward, but it is a poor mechanism for wholesale replication. Use remote write for that.

Precompute with recording rules

Dashboards that recompute expensive aggregations on every refresh are the main source of query load. Move those expressions into recording rules so the aggregation happens once, on a schedule, and the dashboard reads a cheap series.

Standardize dashboards as code, provisioned from a repository with template variables for cluster, namespace, and workload. One well-maintained set of golden-signal dashboards beats four hundred hand-built ones.

Alert on symptoms and error budgets

Page on user-visible symptoms — availability, latency, saturation — rather than on individual resource thresholds. Use multi-window, multi-burn-rate rules against SLO error budgets so a fast burn pages immediately and a slow burn opens a ticket. Route non-urgent conditions to a queue, never to a pager.

Every alert needs a runbook link and an owner. Alerts that nobody can act on are the ones that train people to ignore the pager.

Key takeaways

  • Set an explicit cardinality budget and drop high-churn labels at scrape time.
  • Run per-cluster Prometheus with remote write into a long-term store.
  • Precompute dashboard aggregations with recording rules.
  • Provision dashboards as code with cluster and namespace variables.
  • Alert on SLO burn rate and user-visible symptoms, each with a runbook.

Need a pod that already works this way?

DevGrid Staffing assembles managed DevOps, platform, and SRE pods with the compliance and delivery practices described here built in from week one.