Readiness and liveness: what each protects

Reading · 6 min · Module 4, lesson 1 of 543 min left in this module

Module 4 · Health checksLesson 1 of 5

Goal: Explain what readiness and liveness each protect, and how one ComputeSphere health check does both jobs.

Key idea

A health check answers two different questions. Readiness: should this spherelet get traffic right now? Liveness: is it stuck, so that starting it again would help? On ComputeSphere, one health check answers both, on two clocks.

Readiness protects your users

A spherelet that fails its check Failure threshold times in a row (3 unless you change it) is taken out of traffic. Requests go to the others. As soon as one check passes, it's back in.

Nothing is restarted. That's the point: a spherelet that's briefly overloaded, or still warming up, just gets a rest.

Readiness is also what makes an update safe. A new version gets traffic only after its first passing check, so the old one keeps serving until then (lesson 5.7.2).

Liveness protects the app from itself

Some failures don't clear on their own: a request that never finishes and blocks everything behind it, or a process that stopped answering but didn't exit. For those, a fresh start is the fix.

So a spherelet that fails 10 checks in a row is restarted. Failures in the first 30 seconds don't count, or during your Initial delay if that's longer. The exact timings are in lesson 5.7.3.

One check, two jobs

Both clocks read the same answer from the same endpoint. That's simpler than tuning two checks, but it means anything that makes your check fail for about ten intervals gets the spherelet restarted, whether or not a restart can help.

Two cases where it can't:

  • The app is still starting. A restart sends it back to the beginning, and the loop repeats. Lesson 6.4.3 covers this one.
  • Something outside the app is down, like a database the check calls. Every spherelet fails together, leaves traffic together, and is restarted together. The database is still down afterwards.

So write the check to fail only for problems that taking the spherelet out of traffic, or starting it again, can actually fix. The next lesson shows how.

And without a health check?

A web service with no health check configured is still checked: ComputeSphere tries to connect to its Port. It gets traffic once the port accepts connections, and is restarted after about two and a half minutes of not accepting them. That catches a crashed app, not one that answers every request with an error.

Background workers have no port, so neither check runs. A worker counts as healthy while its process runs.

Check yourself

A spherelet fails three checks in a row during a traffic spike, then passes the fourth. What happened to it?
Which problem does a liveness restart actually fix?

In the docs