Learning paths / Operate your application / Observability basics

The few signals worth watching

Reading · 6 min · Module 1, lesson 2 of 419 min left in this module

Module 1 · Observability basicsLesson 2 of 4

Goal: Name latency, traffic, errors and saturation, and find each one for a ComputeSphere service.

Key idea

Four numbers tell you whether a service is healthy: latency (how long requests take), traffic (how many there are), errors (how many fail) and saturation (how close it is to its limits). Watch these four before anything else.

Site reliability engineers call these the four golden signals. Almost every user-facing problem shows up in at least one of them.

Latency: how long requests take

Averages hide pain. If 95 requests take 50 milliseconds (ms) and 5 take 4 seconds, the average looks fine while one user in twenty waits.

So watch the p95: 95% of requests were faster than this, 5% slower. When p95 rises, some of your users are feeling it.

Traffic: how many requests

Requests per minute or hour. Traffic gives the other signals meaning: ten errors in a quiet hour is a different story from ten in a million requests. A sudden drop to zero is a signal too. It often means something in front of your app broke.

Errors: how many fail

The share of requests answered with a 5xx status, the server's own failures (status codes). A 4xx means the caller asked for something wrong, such as a missing page; a few are normal. A jump in 5xx right after a deploy is the classic sign of a bad release.

Saturation: how full it is

How close the service is to the limits of its shape: CPU used against the vCPU it has, memory used against its GB. A Flex spherelet has 0.25 vCPU and 512 MB (Spherelets, shapes and counts).

Saturation warns you before users notice. CPU near its limit slows every request down; memory at its limit ends in a crash. Spotting saturation goes deeper.

Where each one is on ComputeSphere

Open a service in the console:

  • Traffic tab: Requests and Requests over time (traffic), p95 latency (latency), and Status breakdown, the requests grouped by status (errors).
  • Metrics tab: CPU and Memory (saturation), plus Spherelets, how many are running.

The Traffic tab is on web services and private services only. A background worker or cron job serves no requests, so for those you watch its logs for errors and the Metrics tab for saturation.

Check yourself

Your p95 latency doubled after a deploy, but the average barely moved. What does that tell you?
Which signal warns you before users notice anything?

In the docs