Learning paths / Deploy on ComputeSphere / The deployment lifecycle

Health checks: path, port, delays and thresholds

Reading · 6 min · Module 7, lesson 3 of 642 min left in this module

Module 7 · The deployment lifecycleLesson 3 of 6

Goal: Configure a health check that reflects when your app is really ready to serve, with timings that suit how it starts.

3:57 · captions and chapters · narrated with an AI-generated voice
Transcript

Narration uses an AI-generated voice.

[00:00] Where we're going

By the end of this video, you'll set up a health check that says yes only when your app can really serve, with timings that fit how it starts.

[00:10] What gets checked

A health check is a request ComputeSphere sends each spherelet, again and again. Can you take traffic now? For a web service, it's an HTTP request, like GET slash healthz, on port eighty eighty. Any answer from two hundred to three ninety nine passes. Anything else fails. So does no answer in time. Without one, ComputeSphere only checks that your app accepts connections. That can't tell a working app from a broken one.

[00:42] Set it up

Let's add one to the shop sample. It's a slow starter: it needs about a minute before it can serve. In the service's Settings, find Health Check. It's not configured yet, so choose Enable Health Checks.

Endpoint Path is slash healthz. Port is eighty eighty, the service's port. Initial Delay, five seconds. Check Interval, ten. Timeout can stay empty, for three seconds. And Failure Threshold, three. Then choose Update.

The console asks, Redeploy to apply changes? Choose Later. That saves the check for your next redeploy. And here it is, saved. Now choose Redeploy, at the top of the service, and confirm.

[01:31] Watch it work

The service shows Redeploying. A new spherelet starts, next to the old one. Its runtime log shows the checks, every ten seconds. Each one fails with a five oh three, because the app is still starting. About a minute in, it's ready, and its checks pass. Only then does the old spherelet stop. The health check gated the update. And the service is Running.

[01:58] Two clocks

Failing checks do two different things. Think of them as two clocks. The first takes a spherelet out of traffic, after Failure Threshold failures in a row. It goes back in as soon as a check passes. The second restarts it, after ten failures in a row. Those don't count during the first thirty seconds, or during your Initial Delay, if it's longer.

Now take our settings, for an app that isn't ready yet. Checks at five, fifteen and twenty five seconds fail. So at about twenty five seconds, it's out of traffic. That's fine. It had none yet. Restart counting starts at thirty seconds. Ten failures, ten seconds apart, take ninety more. So at about two minutes, it's restarted.

[02:47] Fit the delay

So an app that needs more than about two minutes to start is restarted before it's ever ready. Our shop needs about a minute, so it made it. If yours needs longer, raise Initial Delay to cover its usual start-up. With ninety, the restart comes at about three minutes.

[03:06] Shallow or deep

One more choice: what the check tests. A shallow check answers two hundred once the app has started. A deep check also queries the database. But when that database has a bad minute, every spherelet fails at once. That's a full outage. A good default is shallow, but honest. Say yes only once start-up has really finished.

[03:31] Recap

So, a quick recap. A health check asks each spherelet, can you take traffic now? Failure Threshold failures take it out of traffic. Ten in a row restart it. And Initial Delay should cover your app's usual start-up. Next up, deploy history, and rolling back. I'll see you there.

Key idea

A health check is a request ComputeSphere sends each spherelet, again and again, asking “can you take traffic now?” It gates every update and pulls a failing spherelet out of traffic, so make it answer yes only when the app can really serve.

What gets checked

A web service's health check is an HTTP request to your app, for example GET /healthz on port 8080. Any response from 200 to 399 passes. Anything else, or no answer in time, fails.

Without one, ComputeSphere only checks that your app accepts connections on its port. That can't tell a working app from one that answers every request with an error.

Set it up

  1. Open the service, go to Settings, and find the Health check card. It says Health checks not configured until you choose Add health check.
  2. Fill in Endpoint path (like /healthz), Port (normally your service's port), Initial delay and Check interval, in seconds. Save stays greyed out until all four are filled in.
  3. Choose Save. The console asks Redeploy to apply changes? Choose Redeploy now, or Later to apply it on your next redeploy.
Every field, its limits, and the API names
Console fieldAllowed valuesAPI field
Endpoint pathRequiredpath
PortRequired, 1–65535port
Initial delay (seconds)Required, 0–300initial_delay_seconds
Check interval (seconds)Required, 1–300period_seconds
Timeout (seconds)Optional, 1–60; 3 if emptytimeout_seconds
Failure thresholdOptional, 1–20; 3 if emptyreadiness_failure_threshold

The API sets these on a service's health_check object, with enabled. Updating it redeploys the service unless you also send "skip_redeploy": true. Configured checks are always HTTP; there's no check type to choose.

Timing: two clocks

Failing checks do two different things:

  • Out of traffic after Failure threshold failures in a row. The spherelet goes back in as soon as a check passes.
  • Restarted after 10 failures in a row. Those aren't counted during the first 30 seconds, or during your Initial delay if it's longer.

Worked through with Initial delay 5, Check interval 10 and Failure threshold 3, for an app that isn't ready yet:

  • About 25 seconds after starting (checks at 5, 15 and 25), it's out of traffic. That's fine: it had none yet.
  • About two minutes after starting, it's restarted: counting starts at 30 seconds, and 10 failures 10 seconds apart take 90 more.

So an app that needs more than about two minutes to start is restarted before it's ever ready, over and over. Raise Initial delay to cover its usual start-up time. With 90, the restart comes at about three minutes.

Rules of thumb: an interval of 10 suits most apps; keep the timeout well above how long the check takes under load; a threshold of 3 rides out one slow answer.

Check yourself

Your app takes three minutes to start, and a deploy keeps restarting it. What do you change?

Shallow or deep?

A shallow check (/healthz returns 200 once the app has started and loaded its config) is fast and depends on nothing outside the app. It won't notice an unreachable database.

A deep check also queries the database or an upstream API. It catches a release that can't reach what it needs. But when a shared dependency has a bad minute, every spherelet fails at once, and a slow database becomes a full outage.

A good default: shallow but honest. Return success only after start-up has finished, including a one-time check against the database if the app can't serve without it. Report dependency health in logs or a separate status endpoint instead.

In the docs