Key idea
A health check is a request ComputeSphere sends each spherelet, again and again, asking “can you take traffic now?” It gates every update and pulls a failing spherelet out of traffic, so make it answer yes only when the app can really serve.
What gets checked
A web service's health check is an HTTP request to your app, for example GET /healthz on port 8080. Any response from 200 to 399 passes. Anything else, or no answer in time, fails.
Without one, ComputeSphere only checks that your app accepts connections on its port. That can't tell a working app from one that answers every request with an error.
Set it up
- Open the service, go to Settings, and find the Health check card. It says Health checks not configured until you choose Add health check.
- Fill in Endpoint path (like
/healthz), Port (normally your service's port), Initial delay and Check interval, in seconds. Save stays greyed out until all four are filled in. - Choose Save. The console asks Redeploy to apply changes? Choose Redeploy now, or Later to apply it on your next redeploy.
Every field, its limits, and the API names
| Console field | Allowed values | API field |
|---|---|---|
| Endpoint path | Required | path |
| Port | Required, 1–65535 | port |
| Initial delay (seconds) | Required, 0–300 | initial_delay_seconds |
| Check interval (seconds) | Required, 1–300 | period_seconds |
| Timeout (seconds) | Optional, 1–60; 3 if empty | timeout_seconds |
| Failure threshold | Optional, 1–20; 3 if empty | readiness_failure_threshold |
The API sets these on a service's health_check object, with enabled. Updating it redeploys the service unless you also send "skip_redeploy": true. Configured checks are always HTTP; there's no check type to choose.
Timing: two clocks
Failing checks do two different things:
- Out of traffic after Failure threshold failures in a row. The spherelet goes back in as soon as a check passes.
- Restarted after 10 failures in a row. Those aren't counted during the first 30 seconds, or during your Initial delay if it's longer.
Worked through with Initial delay 5, Check interval 10 and Failure threshold 3, for an app that isn't ready yet:
- About 25 seconds after starting (checks at 5, 15 and 25), it's out of traffic. That's fine: it had none yet.
- About two minutes after starting, it's restarted: counting starts at 30 seconds, and 10 failures 10 seconds apart take 90 more.
So an app that needs more than about two minutes to start is restarted before it's ever ready, over and over. Raise Initial delay to cover its usual start-up time. With 90, the restart comes at about three minutes.
Rules of thumb: an interval of 10 suits most apps; keep the timeout well above how long the check takes under load; a threshold of 3 rides out one slow answer.
Check yourself
Shallow or deep?
A shallow check (/healthz returns 200 once the app has started and loaded its config) is fast and depends on nothing outside the app. It won't notice an unreachable database.
A deep check also queries the database or an upstream API. It catches a release that can't reach what it needs. But when a shared dependency has a bad minute, every spherelet fails at once, and a slow database becomes a full outage.
A good default: shallow but honest. Return success only after start-up has finished, including a one-time check against the database if the app can't serve without it. Report dependency health in logs or a separate status endpoint instead.