Spotting saturation

Reading · 6 min · Module 3, lesson 2 of 431 min left in this module

Module 3 · MetricsLesson 2 of 4

Goal: Tell a busy service from a starved one by the shape of its CPU and memory charts.

Key idea

A busy service uses a lot but still has room. A starved one has hit its ceiling and something gives: CPU at the ceiling makes requests slower, and memory at the ceiling gets the spherelet stopped.

Busy is fine

A CPU line that rises with traffic and falls when it's quiet is a service doing its job. So is memory that steps up once, as a cache fills, and then stays flat. Neither needs action while it stays clear of the ceiling.

CPU at the ceiling: slow, not stopped

A spherelet can't use more CPU than its shape. When the app wants more, it waits its turn, and every request that needs CPU takes longer.

On the chart, the line climbs to the ceiling and goes flat there: 0.25 vCPU on Flex, 1 on Standard. A flat top is the signature. Busy lines wobble; starved lines stop at a limit.

Users feel it as latency. Requests that do real work slow first; cheap ones, like returning a small list, barely change. That's why p95 latency on the Traffic tab is the signal to watch alongside CPU (lesson 6.1.2).

Memory at the ceiling: stopped and started again

Memory can't be made to wait. When an app asks for more than its shape has, the spherelet is stopped and started fresh. The deploy log shows Out of memory: “The service exceeded its memory allotment and was stopped. Consider a larger shape.”

On the chart that looks like a sawtooth: a climb, a sudden drop to where the app starts, then another climb. Each drop is a restart, and anything the app held in memory is gone.

What to do about it

  • CPU flat at the ceiling: more spherelets share the load; a bigger shape makes each request faster. Module 7 covers both, and lesson 5.2.3 has the rule of thumb.
  • Memory sawtooth: a bigger shape buys room. If memory climbs steadily for hours whatever the traffic, it's probably a leak, and a bigger shape only delays the next drop.
  • Close to either ceiling most of the day: act before a spike does it for you.

Check yourself

Your service's CPU chart sits flat at 0.25 vCPU for an hour on Flex, and users report slow pages. What's happening?
The Memory chart climbs for two hours, drops to almost nothing, then climbs again. What will the deploy log show at each drop?