Key idea
When a service breaks, don't guess. Look in the same order every time, and stop at the first place that explains the symptom.
The order
- Status. Running, Failed, or stuck mid-action? It tells you whether the problem is starting the app or running it (lesson 5.1.3).
- Deploy log. Where the rollout stopped, with a label such as Restarting repeatedly or Unhealthy (lesson 6.2.1).
- Runtime log. What your app said last. A crash nearly always prints its reason just before it stops.
- Metrics. Memory at the shape's ceiling, or CPU pinned, on the Metrics tab (lesson 6.3.2).
- Config. The service's Settings: port, variables, secrets, health check. Compare them with what the app expects.
- Recent changes. What's different from the last version that was Running? The Roll back dialog lists versions with their dates (lesson 5.7.4).
Cheap checks come first. The first three take seconds; config and history take reading.
For step 1, this is where each status sits on a deployment's path:
The path of a deployment
Why an order
A fixed order stops you jumping to the cause you saw last time. And if the answer isn't in one place, you've ruled that place out for good.
If users are affected and the last good version is known, roll back first and troubleshoot after.
Many services misbehaving at once?
Open SphereOps in the console's main menu. It shows CPU, memory and spherelets across your environments in one place, so you can see whether one service is struggling or all of them are.
Check yourself