Key idea
Give the platform and people different endpoints. /healthz is for ComputeSphere: 503 until the app has started, then 200 for as long as this spherelet can serve. /status is for you: what the app depends on, and how each of those is doing.
503 until it has started
The shop sample from this module does exactly this. Called during start-up:
GET /healthz
503 {"ready_in_seconds":17,"status":"starting"}
And once it's ready:
GET /healthz
200 {"status":"ok"}
A 200 during start-up would get the spherelet traffic it can't serve yet. A 503 keeps it out until it can, and during an update the old version keeps serving meanwhile.
The pattern
A flag set at the end of start-up, and two routes. In Node:
let ready = false;
async function start() {
await loadCatalog(); // whatever start-up needs
await connectToDatabaseOnce(); // fail start-up if it can't connect
ready = true;
}
app.get("/healthz", (_req, res) => {
if (!ready) return res.status(503).json({ status: "starting" });
res.json({ status: "ok" });
});
app.get("/status", async (_req, res) => {
const database = (await pingDatabase()) ? "connected" : "unreachable";
res.json({ ready, database, version });
});
/healthz does no outside work at all, so it's fast and can't time out waiting for something else.
It still catches a stuck app. The check is answered by the same server that answers your users, so if that server hangs, the check times out too, and after ten misses the spherelet is restarted. That's the one failure a restart really fixes.
Dependencies: once at start, then on /status
Lesson 3.8.1 checked the database inside /healthz. That's a good way to learn the idea, but on ComputeSphere the check also decides restarts (lesson 6.4.1). A database blip would take every spherelet out of traffic and then restart them all, which fixes nothing.
So split it:
- At start-up, if the app can't serve anything without the database, wait for it and stay at 503 until it connects. A release with a wrong connection string then never gets traffic.
- After that,
/statusreports it. The shop's/statusalways answers 200, with fields like"ready":trueand"uptime_seconds":25, so a dependency problem shows up without failing the check. - Real requests that need the database fail with an error, which shows in your logs and error rate.
Still true from Path 3
No authentication on /healthz, a short timeout on anything it does, and nothing secret in either response: say unreachable, never a hostname or a connection string.
Check yourself