Learning paths / Operate your application / Observability basics

Logs, metrics and traces: which answers what

Reading · 7 min · Module 1, lesson 1 of 426 min left in this module

Module 1 · Observability basicsLesson 1 of 4

Goal: Pick the right signal for a question about a running service.

3:19 · captions and chapters · narrated with an AI-generated voice
Transcript

Narration uses an AI-generated voice.

[00:00] Where we're going

Something's off with your service, and you have a question. By the end of this video, you'll know where to look for the answer. Logs, metrics, or traces. Each one answers a different kind of question.

[00:13] Logs

First, logs. A log is a line your app writes while it runs. Order created. Payment refused. Can't connect to the database. Each line is about one moment, so logs are detailed and exact. They answer questions like, why did this request fail? But they're poor at showing a trend. You'd have to count thousands of lines.

[00:37] Metrics

Second, metrics. A metric is a number, measured again and again. CPU used. Memory held. Requests per minute. How long responses take. Plotted on a graph, a metric shows a trend at a glance. Since when? Is it getting worse? Metrics answer that. What they rarely say is why. For that, you go back to the logs.

[01:05] Traces

Third, traces. A trace follows one request through every service it passes. Say the shop calls a payments service, which calls a bank's API. The trace shows each hop, and how long it took. So it answers, which service made this request slow? With just one service, logs and metrics usually tell you enough.

[01:28] Pick by the question

So which one do you open? Start from your question. Why did it crash? Start with the logs. Did the last deploy make things worse? Metrics. Then the logs from just after the deploy. Which of my services is slow? That's a trace. Is it busy, or broken? Metrics, then the logs, to see the errors. Notice the pattern. Metrics first, to find when. Then logs, to find why.

[01:58] The Logs tab

On ComputeSphere, every service gets logs and metrics, with nothing to install. Here's a small shop API on a demo account, with a little traffic going through it.

The Logs tab shows what the app wrote, one line per request. The route, the status, and how long it took. And here's a warning. Someone asked for a product that doesn't exist, and got a four oh four. That's what happened, exactly, at one moment.

[02:28] Metrics tab and tracing

The Metrics tab turns the same service into numbers over time. CPU. Memory. And how many spherelets are running. This is where you'd spot when something changed. Then you'd read the logs from that moment.

What about traces? There's no built-in tracing. If you need traces, your app sends them to a tracing tool you run or subscribe to. OpenTelemetry is the common standard for that.

[02:55] Recap

So, start from your question. Logs say what happened. Metrics say how much, and how fast. Traces follow one request across services. And most of the time, it's metrics first to find when, then logs to find why. Next, the few signals worth watching.

Key idea

Logs say what happened, one event at a time. Metrics say how much and how fast, as numbers over time. Traces follow one request through every service it touches. Start from your question, then pick the signal that answers it.

Logs: what happened

A log is a line your app writes while it runs: "order created", "payment refused", "can't connect to the database". Each line is about one moment, so logs are detailed and exact.

They answer questions like why did this request fail? and what did the app say just before it crashed? They're poor at is this getting worse?, because you'd have to count thousands of lines to see a trend.

Metrics: how much, how fast

A metric is a number measured again and again: CPU used, memory held, requests per minute, how long responses take. Plotted on a graph, a metric shows a trend at a glance.

Metrics answer since when? and is it getting worse?: memory creeping up all week, errors that jumped after the 14:05 deploy. They rarely say why. For that you go back to the logs around the moment the graph changed.

Traces: where the time went

A trace follows one request through every service it passes. If the shop calls a payments service, which calls a bank's API, the trace shows each hop and how long it took.

Traces answer which service made this request slow? They matter most when a request crosses several services. With one service, logs and metrics usually tell you enough.

Pick by the question

Your questionStart withThen
Why did it crash?LogsMetrics, if it ran out of memory
Did the last deploy make things worse?MetricsLogs from just after the deploy
Which of my services is slow?TracesLogs of the slow service
Is it busy or broken?MetricsLogs, to see the errors

The usual move is metrics first to find when, then logs to find why.

A lighter substitute: give every request an ID and write it in each log line. Structured logging in your app shows how.

Check yourself

Errors on your service started climbing sometime this morning. What do you open first?
Your app is a single service with no calls to other services. Do you need tracing to debug it?