Learning paths / Operate your application / Monitoring and alerts

What deserves an alert

Reading · 6 min · Module 6, lesson 1 of 543 min left in this module

Module 6 · Monitoring and alertsLesson 1 of 5

Goal: Decide which conditions should interrupt a person and which belong on a dashboard.

Key idea

An alert interrupts a person, so it should only fire when someone needs to act now. Everything else you want to know about belongs on a dashboard, where you look when you choose to.

Alerts cost attention

Each alert asks someone to stop what they're doing. If most of them turn out to need nothing, people learn to ignore them, and then they miss the one that mattered. That's alert fatigue, and it's how teams end up with alerts nobody reads.

So the question for every alert is: when this fires, what does the person do? If the answer is "look, and usually nothing", it isn't an alert.

Alert on symptoms users feel

Users feel a few things: errors, slowness, and the site being down. Those are symptoms, and they're worth an interruption because someone is being hurt right now.

Causes are different. High CPU, a restart, memory creeping up: they might lead to a symptom, or might not. A service at 90% CPU that still answers quickly is doing its job.

Good alerts, roughly:

  • Errors: the share of requests failing is well above normal.
  • Latency: p95 has been slow for several minutes, not one spike.
  • Saturation that's about to become a symptom: memory close to the shape's limit for a sustained period, because the next step is Out of memory and a restart.

Dashboards for everything else

The rest goes where you look during a review or an investigation: the Metrics and Traffic tabs from lesson 6.1.2, and the logs. A dashboard can show fifty lines; an alert should mean one thing.

Make alerts hard to trigger by accident

  • Sustained, not instant. A metric that crosses a line for ten seconds is noise. Require it to stay there for minutes.
  • Leave room. Set the threshold where action is needed, not where things first look unusual.
  • Severity that means something. High: act now. Low: look today. If everything is high, nothing is.

The next lesson turns this into ComputeSphere alert rules.

Check yourself

Which of these deserves an alert that interrupts someone?

In the docs