Lab 6.L2: a health check that fits the app

Verified lab · 20 min · Module 4, lesson 4 of 525 min left in this module

Module 4 · Health checksLesson 4 of 5

Self-checked for now. Automatic checking arrives with sign-in; until then, tick off each item under Check your work yourself.

Goal: Tune a health check for an app that's slow to start, so it stops being restarted before it's ready.

You need csph pointed at your learn project and lab environment (lesson 5.1.1), room for one Flex spherelet, and curl.

About the sample

START_DELAY: "150" makes the shop take 150 seconds to start, like an app loading a big catalog. Until then /healthz answers 503. REQUIRE_SIGNING_KEY: "false" lets it run without the secret from lab 6.2.4.

On the free trial, stop or delete a service other than shop-api first if two are running.

  1. Deploy the slow shop

    cat > shop-slow.yaml <<'EOF'
    version: "2"
    services:
      - name: shop-slow
        type: web-service
        image: quay.io/computesphere/learn-shop-api:1.4.0
        port: 8080
        plan: flex
        env_vars:
          START_DELAY: "150"
          REQUIRE_SIGNING_KEY: "false"
    EOF
    csph deploy --file shop-slow.yaml
    

    With no health check yet, ComputeSphere only checks that the port answers, so this goes Running early.

    You should seecsph prints Running and the URL within a minute or so.

  2. Measure its start-up

    curl https://<your-shop-slow-url>/healthz
    

    Wait a minute after Running before the first call. If curl says it could not resolve host, wait and retry. Repeat until it says ok: that's the start time your check has to fit.

    You should see{"ready_in_seconds":…,"status":"starting"} at first; about 150 seconds after the start, {"status":"ok"}.

  3. Add a health check that's too eager

    In the console, open shop-slow, Settings. Under Health check, choose Add health check and set Endpoint path /healthz, Port 8080, Initial delay 5, Check interval 10. Choose Save, then Redeploy now, and open Logs, Deploy.

    You should seeCreated and Pulling image in the deploy log: a new version is starting.

  4. Watch the restart loop

    Open Logs, then Deploy, then Runtime. Give it five minutes. Your URL keeps answering: the old version serves until the new one passes. Left alone, the update ends Failed after about ten minutes.

    Check yourself

    Why does the new version never get ready?

    You should seeUnhealthy: Health checks are failing. in the deploy log, then Stopping, and the start-up lines repeating in the runtime log about every two minutes.

  5. Give it time to start

    In Health check, choose Edit, set Initial delay to 170, then Save and Redeploy now.

    You should seeLive in the deploy log about three minutes later, and no more restarts.

  6. Check it answers

    curl https://<your-shop-slow-url>/healthz
    

    You should see{"status":"ok"}

Check your work

  • If this doesn't pass

    Open Settings, Health check. If it still shows 5s, choose Edit, change Initial delay, then Save and Redeploy now.

  • If this doesn't pass

    No Live in the deploy log after five minutes? Unhealthy there means the new version still has the old Initial delay: change it again and choose Redeploy now.

  • If this doesn't pass

    A 503 with ready_in_seconds means a spherelet has just started. Wait for the count to reach zero and try again.

Clean up
csph services list
csph services delete <shop-slow-service-id>
rm shop-slow.yaml