Key idea
Logs say what happened, one event at a time. Metrics say how much and how fast, as numbers over time. Traces follow one request through every service it touches. Start from your question, then pick the signal that answers it.
Logs: what happened
A log is a line your app writes while it runs: "order created", "payment refused", "can't connect to the database". Each line is about one moment, so logs are detailed and exact.
They answer questions like why did this request fail? and what did the app say just before it crashed? They're poor at is this getting worse?, because you'd have to count thousands of lines to see a trend.
Metrics: how much, how fast
A metric is a number measured again and again: CPU used, memory held, requests per minute, how long responses take. Plotted on a graph, a metric shows a trend at a glance.
Metrics answer since when? and is it getting worse?: memory creeping up all week, errors that jumped after the 14:05 deploy. They rarely say why. For that you go back to the logs around the moment the graph changed.
Traces: where the time went
A trace follows one request through every service it passes. If the shop calls a payments service, which calls a bank's API, the trace shows each hop and how long it took.
Traces answer which service made this request slow? They matter most when a request crosses several services. With one service, logs and metrics usually tell you enough.
Pick by the question
| Your question | Start with | Then |
|---|---|---|
| Why did it crash? | Logs | Metrics, if it ran out of memory |
| Did the last deploy make things worse? | Metrics | Logs from just after the deploy |
| Which of my services is slow? | Traces | Logs of the slow service |
| Is it busy or broken? | Metrics | Logs, to see the errors |
The usual move is metrics first to find when, then logs to find why.
A lighter substitute: give every request an ID and write it in each log line. Structured logging in your app shows how.
Check yourself