Learning paths / Operate your application / Troubleshooting playbook

Try it: five broken deploys

Reading · 30 min · Module 5, lesson 3 of 435 min left in this module

Module 5 · Troubleshooting playbookLesson 3 of 4

Goal: Deploy five broken services, name each fault from its status and logs, fix it, and clean up.

You need csph signed in, with learn and lab as your defaults (lesson 5.1.1), and curl.

Room on your account

Each case is a new service, shop-case-1 to shop-case-5, so none of them touches your other services. A trial account runs two Flex spherelets in total, so do the cases one at a time and delete each before the next. If a deploy is refused for capacity, stop or delete another service first.

  1. Download the five cases

    mkdir broken-deploys && cd broken-deploys
    curl -fsSLO "https://raw.githubusercontent.com/computesphere-samples/learn/main/labs/shop-api/broken/case-[1-5].yaml"
    ls
    

    Each file deploys the shop sample with one thing wrong. Don't read them yet: find the fault the way you would in production.

    You should seecase-1.yaml to case-5.yaml in a new folder.

  2. Read the evidence for case 1

    csph deploy --file case-1.yaml --no-wait
    

    --no-wait returns at once. Wait about two minutes for it to fail, then read the logs (csph lists the newest deploy event first):

    csph logs shop-case-1 --kind deploy
    csph logs shop-case-1
    

    Follow the order from lesson 6.5.1: status (on its page in the console), deploy log, runtime log, then the config in the file. Name the fault.

    You should seethe deploy log shows Unhealthy, and the runtime log ends with listening.

  3. Fix it with a fresh start, then clean up

    A port change doesn't reach the URL of a service that already exists. So delete the broken service, fix the file, and deploy it again: a new service gets everything in the file.

    csph services list
    csph services delete <shop-case-1-service-id>
    csph services list -o json
    csph deploy --file case-1.yaml
    

    Once it answers, delete shop-case-1 the same way and check the JSON list. Clean up after every case, so the next one has room.

    You should seeshop-case-1 comes back Running and its URL answers; after the last delete, it isn't in the JSON list.

  4. Case 2

    Same evidence loop, with case-2.yaml, --no-wait and shop-case-2. The fix is a secret, set on the running service, as in lesson 5.4.3:

    csph services secrets set <KEY>=demo-not-a-real-key-0000 --service <service-id>
    

    Setting it redeploys the service.

    You should seeRestarting repeatedly in the deploy log, and a fatal line in the runtime log.

  5. Case 3

    Same loop, with case-3.yaml. When you know which variable is to blame, delete its line from the file and deploy again; a variable the file no longer lists is removed:

    csph deploy --file case-3.yaml
    

    You should seeThe runtime log stops partway through starting, and the deploy log names the reason.

  6. Case 4

    csph deploy --file case-4.yaml --no-wait
    

    Read the error. It still leaves an empty shop-case-4 service: fix the file and deploy it again, and the service gets its first version.

    You should seeAn error straight away. Nothing starts, so there are no logs.

  7. Case 5

    Deploy case-5.yaml and look. Fix the file and deploy it again.

    You should seeUnhealthy in the deploy log, and requests answered 404 in the runtime log.

Answers (try first)
  1. Wrong port. The service's port is 3000; the app listens on 8080. Delete the service, set port: 8080, and deploy the file again.
  2. Crash loop. The fatal line names SIGNING_KEY, and the deploy log says Restarting repeatedly. Set the SIGNING_KEY secret; that redeploys it.
  3. Out of memory. CACHE_MB: "900" asks for 900 MB on a 512 MB Flex spherelet. The log stops after warming the product cache. Delete the CACHE_MB line and deploy the file again.
  4. Image can't be used. Tag 1.0.9 doesn't exist. Use the tag the other cases use, and deploy the file again.
  5. Bad config. The health check is on /health, which answers 404. Set health_check_path: /healthz and deploy the file again.

Check yourself

Which case left no logs at all, and why?

In the docs