Learning paths / Containers / Containerize a real app

A health endpoint in every app

Reading · 6 min · Module 8, lesson 1 of 659 min left in this module

Module 8 · Containerize a real appLesson 1 of 6

Goal: Add a /healthz endpoint that answers 200 only when the app can really serve, checking the dependencies it can't work without.

Key idea

A health endpoint, usually /healthz, is a route that answers one question for the platform: "can you take traffic right now?" It returns 200 when the answer is yes and an error status such as 503 when it's no. Every app you containerize should have one.

Why a container needs one

A running container only proves the process started. It doesn't prove the app finished starting, that it can reach its database, or that it isn't stuck. A platform can't tell those apart unless the app says so.

With a health endpoint, the platform calls it again and again. A new version gets traffic only once it answers 200, and a copy that starts failing is taken out of traffic. Path 5 shows how ComputeSphere does this in lesson 5.7.3, Health checks.

The simplest version

The three samples in this module each have one. In Node:

app.get("/healthz", (_req, res) => {
  res.json({ status: "ok" });
});

If the process can run this route, it's healthy. For an app with no dependencies, that's the whole answer. It also catches more than you'd think: an app that's still starting, or one whose server has stopped answering, never gets to send the 200.

Checking what the app needs

An app that can't serve a single page without its database isn't healthy when the database is down. So check it, quickly. This is the web app from the Compose lessons, trimmed:

app.get("/healthz", async (_req, res) => {
  const body = { status: "ok" };
  if (databaseUrl) body.database = (await checkDatabase()) ? "connected" : "unreachable";

  const healthy = body.database !== "unreachable";
  if (!healthy) body.status = "unhealthy";
  res.status(healthy ? 200 : 503).json(body);
});

checkDatabase() connects with a 2-second timeout and runs SELECT 1. With the database unreachable, the endpoint answers:

HTTP/1.1 503 Service Unavailable
{"status":"unhealthy","database":"unreachable"}

Rules that keep it useful

  • Be fast. Health checks time out after a few seconds. Give every dependency check a short timeout of its own.
  • Check only what you can't serve without. A payment provider or email service being slow shouldn't take your whole app out of traffic. Handle those failures in the routes that use them.
  • Skip authentication. The platform calls it without credentials.
  • Reveal nothing. Say connected or unreachable, never a connection string or an error with a hostname in it.
  • Keep it cheap. It runs over and over, on every copy, so no heavy queries.

Check yourself

Your app sends emails through a third-party provider. Should /healthz fail when the provider is down?
Your database check has no timeout, and the database is hanging. What does the platform see?

In the docs