Key idea
A service stays up when there's always a healthy copy to answer. Rolling updates give you that during a deploy, even on one spherelet. Only a second spherelet gives you that when a copy crashes.
Through a deploy
Lesson 5.7.2 showed the rule: a new spherelet takes traffic only after it passes its health check, and an old one retires only after that. With two spherelets, they're replaced one at a time, and at least two are running throughout. On ComputeSphere today, about one request can still fail (a 503 or a timeout) as each old spherelet retires, so have clients retry safe requests.
Zero-downtime update · old serves until new is healthy
A new spherelet starts alongside the old ones. It gets no traffic yet.
It gets traffic, and one old spherelet retires, with up to 30 seconds to finish its requests.
Every spherelet runs the new version. Never fewer than 2 serving.
- The old spherelets keep serving all the traffic.
- The new one is restarted while its checks keep failing.
- After about ten minutes the deployment is marked Failed. The previous version is still live on the same URL.
This works on one spherelet too, so a deploy alone isn't a reason to run two. Two things make it hold:
- A health check that means ready (module 4). Without one, a new spherelet counts as ready as soon as its port accepts connections, and a broken release can go live.
- A clean shutdown. Finish in-flight requests when the stop signal arrives, then exit.
Through a crash
This is where the count matters.
- On one spherelet, a crash is an outage. Nothing answers until ComputeSphere has restarted it and it passes its health check again, which for a slow starter can be minutes.
- On two, the health check takes the failing spherelet out of traffic and the other keeps answering while it restarts (module 4, lesson 6.4.1). Users see, at most, the few requests that were on the crashed copy.
The same goes for maintenance on the platform. When ComputeSphere moves spherelets onto new hardware, a service with two or more is moved one at a time and keeps answering. A single spherelet can have a short gap while its replacement starts.
What two spherelets ask of you
- The app must be stateless (lesson 6.7.2). A volume attaches to one spherelet at a time, so a service with a volume stays at one; its availability comes from quick restarts and good backups instead.
- It costs twice as much. On autoscaling, set the minimum to 2 for the same effect.
- Watch the crash, not just the uptime. With two, a spherelet crash-looping every hour can go unnoticed, because the other covers for it. Check the deploy log for Restarting repeatedly now and then.
Check yourself