Learning paths / Operate your application / Backups and availability

Availability: more than one spherelet, rolling updates

Reading · 6 min · Module 8, lesson 2 of 543 min left in this module

Module 8 · Backups and availabilityLesson 2 of 5

Goal: Keep a service answering through a deploy and through one spherelet crashing.

2:56 · captions and chapters · narrated with an AI-generated voice
Transcript

Narration uses an AI-generated voice.

[00:00] Where we're going

By the end of this video, you'll watch a service keep answering while it's updated, one spherelet at a time.

[00:06] One or two

A service stays up when there's always a healthy copy to answer. Take one spherelet. If it crashes, nothing answers. Users see errors until it has restarted, and passed its health check again. Now take two. The health check takes the failing one out of traffic. And the other keeps answering while it restarts.

[00:31] Through a deploy

A deploy is different. Rolling updates cover it. First, one new spherelet starts beside the old ones. It gets no traffic yet. When its health check passes, it takes traffic, and one old spherelet retires. Then the same again, until every spherelet runs the new version. This works on one spherelet too. So it's a crash, not a deploy, that calls for two.

[00:57] Watch it happen

Let's watch one. Here's the shop sample, on two Flex spherelets, with a health check. The Health card shows both spherelets, healthy. We'll make a real change: a new shop name. In Settings, under Variables, change the value, and choose Save. Then choose Save and redeploy.

Meanwhile, a loop sends the service a request every half second. It prints the time, the status, and the shop name. A new spherelet joins. That's three now, all healthy. Then an old one retires, and answers start coming from the new version. A second new spherelet joins, and the last old one retires. Two again, both new.

[01:45] Count every answer

So, did it keep answering? Let's count. Nine hundred and eighty three requests. Nine hundred and eighty one answered two hundred. Two got no answer within five seconds. Each came right as an old spherelet retired. That's a gap on ComputeSphere today, and it's been reported. So when you test a deploy, count every answer, the way this loop does.

[02:10] What two ask of you

Two spherelets ask three things of you. First, the app must be stateless. A volume attaches to one spherelet at a time. Second, it costs twice as much. Two Flex spherelets are eighteen dollars a month. Third, watch the crashes, not just the uptime. The second spherelet can hide one that keeps restarting.

[02:34] Recap

So, a quick recap. Rolling updates keep the old spherelet serving until the new one is healthy. A second spherelet covers a crash. And a service on two must be stateless. Next up, snapshots, daily schedules and restores. I'll see you there.

Key idea

A service stays up when there's always a healthy copy to answer. Rolling updates give you that during a deploy, even on one spherelet. Only a second spherelet gives you that when a copy crashes.

Through a deploy

Lesson 5.7.2 showed the rule: a new spherelet takes traffic only after it passes its health check, and an old one retires only after that. With two spherelets, they're replaced one at a time, and at least two are running throughout. On ComputeSphere today, about one request can still fail (a 503 or a timeout) as each old spherelet retires, so have clients retry safe requests.

This works on one spherelet too, so a deploy alone isn't a reason to run two. Two things make it hold:

  • A health check that means ready (module 4). Without one, a new spherelet counts as ready as soon as its port accepts connections, and a broken release can go live.
  • A clean shutdown. Finish in-flight requests when the stop signal arrives, then exit.

Through a crash

This is where the count matters.

  • On one spherelet, a crash is an outage. Nothing answers until ComputeSphere has restarted it and it passes its health check again, which for a slow starter can be minutes.
  • On two, the health check takes the failing spherelet out of traffic and the other keeps answering while it restarts (module 4, lesson 6.4.1). Users see, at most, the few requests that were on the crashed copy.

The same goes for maintenance on the platform. When ComputeSphere moves spherelets onto new hardware, a service with two or more is moved one at a time and keeps answering. A single spherelet can have a short gap while its replacement starts.

What two spherelets ask of you

  • The app must be stateless (lesson 6.7.2). A volume attaches to one spherelet at a time, so a service with a volume stays at one; its availability comes from quick restarts and good backups instead.
  • It costs twice as much. On autoscaling, set the minimum to 2 for the same effect.
  • Watch the crash, not just the uptime. With two, a spherelet crash-looping every hour can go unnoticed, because the other covers for it. Check the deploy log for Restarting repeatedly now and then.

Check yourself

A web service runs on 1 spherelet with a good health check. You deploy a new version. Do users see downtime?
The same service's only spherelet crashes at 3 a.m. What do users see?