Learning paths / Run in production

Learning path 68 modules38 lessons open

Operate your application

Logs, metrics, health checks, troubleshooting, alerts, scaling and backups for apps real people use.

At a glance

Level
Intermediate
Duration
About 7 hours
Modules
8
Verified labs
5
Prerequisites
Deploy on ComputeSphere

Modules in this path

  1. 1. Observability basicsLogs, metrics and traces
    1. Logs, metrics and traces: which answers what Reading, 7 min
    2. The few signals worth watching Reading, 6 min
    3. Try it: a tour of your service's Logs, Metrics and Traffic tabs Reading, 8 min
    4. Check: observability Knowledge check, 5 min
    4 open
  2. 2. LogsBuild, deploy and runtime logs
    1. Build, deploy and runtime logs Reading, 6 min
    2. Structured logging in your app Reading, 7 min
    3. Searching and streaming logs, including cron runs Video and reading, 10 min
    4. Lab: find the cause in the logs Verified lab, 25 min
    5. Check: logs Knowledge check, 5 min
    5 open
  3. 3. MetricsCPU, memory, spherelets, saturation
    1. CPU, memory and spherelet count Reading, 7 min
    2. Spotting saturation Reading, 6 min
    3. Try it: load the sample and read the graphs Reading, 20 min
    4. Check: metrics Knowledge check, 5 min
    4 open
  4. 4. Health checksReadiness, liveness and slow starts
    1. Readiness and liveness: what each protects Reading, 6 min
    2. A good health endpoint Reading, 7 min
    3. Slow starters and the restart loop Reading, 5 min
    4. Lab 6.L2: a health check that fits the app Verified lab, 20 min
    5. Check: health checks Knowledge check, 5 min
    5 open
  5. 5. Troubleshooting playbookA repeatable order of checks
    1. A troubleshooting order that always works Video and reading, 12 min
    2. Failure gallery: wrong port, crash loop, out of memory, image pull, bad config Reading, 7 min
    3. Try it: five broken deploys Reading, 30 min
    4. Check: troubleshooting Knowledge check, 5 min
    4 open
  6. 6. Monitoring and alertsRules, thresholds, notifications
    1. What deserves an alert Reading, 6 min
    2. Alert rules: CPU and memory, threshold, evaluation period, severity Reading, 7 min
    3. Where notifications go: in-app, email and webhooks Reading, 5 min
    4. Lab: create an alert Verified lab, 20 min
    5. Check: alerts Knowledge check, 5 min
    5 open
  7. 7. ScalingBigger shapes, more spherelets, autoscaling
    1. Vertical scaling: bigger shapes Reading, 6 min
    2. Horizontal scaling: more spherelets Reading, 6 min
    3. Autoscaling on CPU: minimum, maximum, target Reading, 7 min
    4. Lab: scale out, then autoscale Verified lab, 25 min
    5. Check: scaling Knowledge check, 5 min
    5 open
  8. 8. Backups and availabilitySnapshots, restores, zero-downtime
    1. Backups, snapshots and standby copies Reading, 7 min
    2. Availability: more than one spherelet, rolling updates Reading, 6 min
    3. Snapshots, daily schedules and restores Reading, 7 min
    4. Lab: back up and restore Verified lab, 25 min
    5. Check: backups and availability Knowledge check, 5 min
    5 open
Path challenge, 60 minChallenge: Troubleshoot a failing deployRestore a broken service on your own, by finding every fault a bad release brought in and fixing it forward.