A good health endpoint

Reading · 7 min · Module 4, lesson 2 of 537 min left in this module

Module 4 · Health checksLesson 2 of 5

Goal: Write a health endpoint that answers 503 until start-up finishes, and report dependencies on a separate status endpoint so the check doesn't become fragile.

Key idea

Give the platform and people different endpoints. /healthz is for ComputeSphere: 503 until the app has started, then 200 for as long as this spherelet can serve. /status is for you: what the app depends on, and how each of those is doing.

503 until it has started

The shop sample from this module does exactly this. Called during start-up:

GET /healthz
503 {"ready_in_seconds":17,"status":"starting"}

And once it's ready:

GET /healthz
200 {"status":"ok"}

A 200 during start-up would get the spherelet traffic it can't serve yet. A 503 keeps it out until it can, and during an update the old version keeps serving meanwhile.

The pattern

A flag set at the end of start-up, and two routes. In Node:

let ready = false;

async function start() {
  await loadCatalog();            // whatever start-up needs
  await connectToDatabaseOnce();  // fail start-up if it can't connect
  ready = true;
}

app.get("/healthz", (_req, res) => {
  if (!ready) return res.status(503).json({ status: "starting" });
  res.json({ status: "ok" });
});

app.get("/status", async (_req, res) => {
  const database = (await pingDatabase()) ? "connected" : "unreachable";
  res.json({ ready, database, version });
});

/healthz does no outside work at all, so it's fast and can't time out waiting for something else.

It still catches a stuck app. The check is answered by the same server that answers your users, so if that server hangs, the check times out too, and after ten misses the spherelet is restarted. That's the one failure a restart really fixes.

Dependencies: once at start, then on /status

Lesson 3.8.1 checked the database inside /healthz. That's a good way to learn the idea, but on ComputeSphere the check also decides restarts (lesson 6.4.1). A database blip would take every spherelet out of traffic and then restart them all, which fixes nothing.

So split it:

  • At start-up, if the app can't serve anything without the database, wait for it and stay at 503 until it connects. A release with a wrong connection string then never gets traffic.
  • After that, /status reports it. The shop's /status always answers 200, with fields like "ready":true and "uptime_seconds":25, so a dependency problem shows up without failing the check.
  • Real requests that need the database fail with an error, which shows in your logs and error rate.

Still true from Path 3

No authentication on /healthz, a short timeout on anything it does, and nothing secret in either response: say unreachable, never a hostname or a connection string.

Check yourself

Your API's database goes down for five minutes. With the pattern above, what happens to your spherelets?
Why should /healthz answer 503, not 200, while the app is still loading?

In the docs