Slow starters and the restart loop

Reading · 5 min · Module 4, lesson 3 of 530 min left in this module

Module 4 · Health checksLesson 3 of 5

Goal: Recognise a restart loop in the deploy and runtime logs, and fix it by raising Initial delay above the app's start-up time.

3:04 · captions and chapters · narrated with an AI-generated voice
Transcript

Narration uses an AI-generated voice.

[00:00] Where we're going

Some apps are healthy, but slow to start. Check them too early, and they're restarted before they're ever ready. By the end of this video, you'll spot that loop in the logs, and fix it with one setting.

[00:13] A slow starter

Here's the shop sample again. This time, it takes a hundred and fifty seconds to start, like an app loading a big catalog. We give it a health check that's too eager. Endpoint path, slash healthz. Initial delay, five seconds. Check interval, ten. Then Save, and Redeploy now.

[00:33] Spot the loop

Open Logs, then Deploy. The new version is created, and its image is pulled. Then, Unhealthy. Health checks are failing. Then, Stopping. Your URL keeps answering, though. The old version serves until the new one passes.

Now the runtime log. Show just the info lines. Starting. Loading the catalog. Two minutes later, stopped. And starting again, from zero. It repeats every two minutes, and the ready line never comes. Notice what's missing. There's no error, and no crash. The app is fine, just slow. The restarts are what stop it.

[01:19] Two clocks

Why two minutes? Failing checks run two clocks. The first takes a spherelet out of traffic, after three failures in a row. Checks at five, fifteen and twenty five seconds. This one had no traffic yet, so that's harmless. The second clock restarts it, after ten failures in a row. Counting starts at thirty seconds, or at the Initial delay, if that's later. Ten checks, ten seconds apart, end at about a hundred and twenty seconds. That's the restart. Our shop needs a hundred and fifty. It never gets there.

[01:55] The fix

So, first, measure the start-up. Here, it's a hundred and fifty seconds. Then set Initial delay a little above it. A hundred and seventy gives some room. Now restart counting starts at a hundred and seventy. The first check there passes, because the shop is ready.

In Health check, choose Edit. Set Initial delay to a hundred and seventy. Save, and Redeploy now. About three minutes later, the deploy log says Live. And the runtime log shows the ready line, just once. Don't just set three hundred everywhere, though. Every deploy waits out the delay before it's ready.

[02:37] Recap

So, a quick recap. Unhealthy in the deploy log, and start-up lines repeating with no error. That's a restart loop. The restart comes ten failures after thirty seconds, or after the Initial delay. So set Initial delay a little above your app's start-up time. Next up, you'll do this yourself, in the lab.

Key idea

An app that takes longer to start than its health check waits is restarted before it's ready. It starts again from zero, is restarted again, and never gets there. The fix is an Initial delay longer than its start-up.

What the loop looks like

  • Status never reaches Running. A new service sits on Deploying. An update sits on Redeploying while the old version keeps serving, and reads Failed after about ten minutes.
  • Deploy log: Unhealthy, “Health checks are failing.”, then Stopping.
  • Runtime log: the start-up lines repeat every couple of minutes, and the “ready” line never comes.

The shop sample logs failing checks, so its runtime log shows the loop plainly:

{"time":"…","level":"info","msg":"starting learn-shop-api","version":"1.0.0",…,"start_delay":150,…}
{"time":"…","level":"info","msg":"loading the catalog; /healthz answers 503 until it's done","start_delay":150}
{"time":"…","level":"warn","msg":"health check failed: still starting","route":"/healthz","status":503,…}
…
{"time":"…","level":"info","msg":"starting learn-shop-api","version":"1.0.0",…,"start_delay":150,…}

A crash looks different: the app exits with an error line, and the deploy log says Restarting repeatedly. Here nothing crashes. The app is healthy, just slow, and the restarts are what stop it.

Why it happens

A restart comes after 10 failed checks in a row, and counting starts at the Initial delay or at 30 seconds, whichever is later. With Initial delay 5 and Check interval 10, the first counted failure is at 30 seconds and the tenth is nine intervals later: 30 + 9 × 10 = about 120 seconds. An app that needs 150 seconds never makes it. With a Check interval of 30, the same app would have until about 300 seconds. Lesson 5.7.3 has every field and its limits.

The fix

  1. Measure the start-up. In the runtime log, compare the timestamp of the first start-up line with the “ready” line from a start that finished (run it locally if it never finishes on ComputeSphere).
  2. Set Initial delay a little above it. For 150 seconds, 170 gives some room. In the console: Settings, Health check, Edit, Save, then Redeploy now. A manifest only sets the Endpoint path, so timings live in the console.

Don't just set 300 everywhere. The first check waits for the Initial delay, so every deploy takes at least that long to reach Running.

Start-up longer than five minutes?

Initial delay tops out at 300 seconds. A longer Check interval pushes the restart further out (it's 9 intervals after the delay), but it also slows how fast a real failure is noticed.

The better fix is a faster start: load only what the first request needs, and fill caches that aren't essential after /healthz starts answering 200.

Check yourself

Initial delay 5, Check interval 10, and the app takes 100 seconds to start. Is it caught in a restart loop?
A new service stays on Deploying. The deploy log repeats Unhealthy, and the runtime log shows the same start-up lines every two minutes with no errors. What do you change?