Key idea
An app that takes longer to start than its health check waits is restarted before it's ready. It starts again from zero, is restarted again, and never gets there. The fix is an Initial delay longer than its start-up.
What the loop looks like
- Status never reaches Running. A new service sits on Deploying. An update sits on Redeploying while the old version keeps serving, and reads Failed after about ten minutes.
- Deploy log: Unhealthy, “Health checks are failing.”, then Stopping.
- Runtime log: the start-up lines repeat every couple of minutes, and the “ready” line never comes.
The shop sample logs failing checks, so its runtime log shows the loop plainly:
{"time":"…","level":"info","msg":"starting learn-shop-api","version":"1.0.0",…,"start_delay":150,…}
{"time":"…","level":"info","msg":"loading the catalog; /healthz answers 503 until it's done","start_delay":150}
{"time":"…","level":"warn","msg":"health check failed: still starting","route":"/healthz","status":503,…}
…
{"time":"…","level":"info","msg":"starting learn-shop-api","version":"1.0.0",…,"start_delay":150,…}
A crash looks different: the app exits with an error line, and the deploy log says Restarting repeatedly. Here nothing crashes. The app is healthy, just slow, and the restarts are what stop it.
Why it happens
A restart comes after 10 failed checks in a row, and counting starts at the Initial delay or at 30 seconds, whichever is later. With Initial delay 5 and Check interval 10, the first counted failure is at 30 seconds and the tenth is nine intervals later: 30 + 9 × 10 = about 120 seconds. An app that needs 150 seconds never makes it. With a Check interval of 30, the same app would have until about 300 seconds. Lesson 5.7.3 has every field and its limits.
The fix
- Measure the start-up. In the runtime log, compare the timestamp of the first start-up line with the “ready” line from a start that finished (run it locally if it never finishes on ComputeSphere).
- Set Initial delay a little above it. For 150 seconds, 170 gives some room. In the console: Settings, Health check, Edit, Save, then Redeploy now. A manifest only sets the Endpoint path, so timings live in the console.
Don't just set 300 everywhere. The first check waits for the Initial delay, so every deploy takes at least that long to reach Running.
Start-up longer than five minutes?
Initial delay tops out at 300 seconds. A longer Check interval pushes the restart further out (it's 9 intervals after the delay), but it also slows how fast a real failure is noticed.
The better fix is a faster start: load only what the first request needs, and fill caches that aren't essential after /healthz starts answering 200.
Check yourself