Learning paths / Operate your application / Troubleshooting playbook

A troubleshooting order that always works

Video and reading · 12 min · Module 5, lesson 1 of 454 min left in this module

Module 5 · Troubleshooting playbookLesson 1 of 4

Goal: Work through a broken service in the same order every time, from its status to what changed recently.

2:41 · captions and chapters · narrated with an AI-generated voice
Transcript

Narration uses an AI-generated voice.

[00:00] Where we're going

When a service breaks, it's tempting to guess. By the end of this video, you'll have an order to follow instead. And we'll follow it on a real broken deploy.

[00:10] The order

There are six places to look, in the same order every time. First, status. Is it Running, Failed, or stuck partway through? Second, the deploy log. It shows where the rollout stopped. Third, the runtime log. That's what your app said last. Fourth, metrics. Is memory or CPU at its ceiling? Fifth, config. The port, the variables and secrets, the health check. And sixth, recent changes. What's different from the last version that worked? Cheap checks come first. The first three take seconds. And here's the one rule. Stop at the first place that explains the problem.

[00:53] A real broken deploy

Here's a real one. The shop API, freshly deployed, and something's wrong. Let's go in order.

Step one, status. Minutes later, it still says Deploying. And Health shows zero of one spherelets healthy. So the problem is starting the app, not running it.

Step two, the deploy log. Open Logs, then Deploy. There's the label. Restarting repeatedly. The app exits while it starts, gets restarted, and exits again.

[01:26] Find the cause

Step three, the runtime log. A crash nearly always prints its reason just before it stops.

And here it is, again and again. A fatal line: the signing key is not set. The shop can't sign orders without it.

That explains the symptom, so we stop looking. Metrics and recent changes can wait. The fix is in config.

[01:51] Fix it

In Settings, the Secrets tab is empty. Nothing sets the signing key.

So we set it as a secret, from the terminal. It's a demo value, not a real key. Setting a secret redeploys the service.

Soon after, it's Running, and its spherelet is healthy. And the runtime log now says listening.

[02:13] Recap

So, when something breaks, don't guess. Status, deploy log, runtime log. Then metrics, config, and recent changes. Stop at the first place that explains it. And if users are affected and you know the last good version, roll back first. Troubleshoot after. Next, you'll see five common failures, and what each one looks like.

Key idea

When a service breaks, don't guess. Look in the same order every time, and stop at the first place that explains the symptom.

The order

  1. Status. Running, Failed, or stuck mid-action? It tells you whether the problem is starting the app or running it (lesson 5.1.3).
  2. Deploy log. Where the rollout stopped, with a label such as Restarting repeatedly or Unhealthy (lesson 6.2.1).
  3. Runtime log. What your app said last. A crash nearly always prints its reason just before it stops.
  4. Metrics. Memory at the shape's ceiling, or CPU pinned, on the Metrics tab (lesson 6.3.2).
  5. Config. The service's Settings: port, variables, secrets, health check. Compare them with what the app expects.
  6. Recent changes. What's different from the last version that was Running? The Roll back dialog lists versions with their dates (lesson 5.7.4).

Cheap checks come first. The first three take seconds; config and history take reading.

For step 1, this is where each status sits on a deployment's path:

Why an order

A fixed order stops you jumping to the cause you saw last time. And if the answer isn't in one place, you've ruled that place out for good.

If users are affected and the last good version is known, roll back first and troubleshoot after.

Many services misbehaving at once?

Open SphereOps in the console's main menu. It shows CPU, memory and spherelets across your environments in one place, so you can see whether one service is struggling or all of them are.

Check yourself

A new version shows Failed. Its runtime log is empty. Where do you look next?