Learning paths / Deploy on ComputeSphere / The deployment lifecycle

Rolling updates: the old version keeps serving

Reading · 5 min · Module 7, lesson 2 of 647 min left in this module

Module 7 · The deployment lifecycleLesson 2 of 6

Goal: Explain why the old version keeps serving until the new one is healthy, and what happens when it never is.

Key idea

During an update, a new spherelet only gets traffic after it passes its health check, and an old one only retires after that. So a bad release never reaches your users: the old version keeps serving.

This happens every time you deploy a new image, redeploy with changes, or restart a web service.

The rule: never fewer than you asked for

ComputeSphere never drops below your spherelet count during an update, and adds at most one extra spherelet at a time. The diagram shows a service on 2 spherelets.

With 1 spherelet it's the same: the new one starts next to the old, and the old one answers until the new one is ready.

One caveat on ComputeSphere today: as each old spherelet retires, a request that reaches it at that moment can fail with a 503 or time out. That's about one request per retired spherelet, not an outage. Clients that retry safe requests, such as a GET, won't notice.

The gate: what “healthy” means

  • With a health check (next lesson), healthy means your endpoint answered successfully.
  • Without one, healthy only means your app accepts connections on its port.

That difference matters. An app that listens on its port but can't reach its database passes the default check and goes live, broken. A health check that tests real readiness stops that release at the gate.

When the new version never becomes healthy

The update can't take its first step. The old spherelets keep serving all traffic, the new one is restarted while its checks fail, and after about ten minutes the deployment is marked Failed.

Failed here means “this release didn't go out”, not “your service is down”. The previous version is still live on the same URL. Deploy a fix or roll back (lesson 5.7.4); the lab at the end of this module does exactly that.

The gate also protects a Running service. A spherelet that starts failing its health check is taken out of traffic until it passes again, and restarted if it keeps failing.

Check yourself

A new release fails its health check for ten minutes. What do your users see during that time?

What your app has to do

  • Answer the health check only when it can serve. Returning success before the config is loaded sends traffic too early.
  • Shut down cleanly. A retired spherelet gets a stop signal. Finish in-flight requests and exit; requests longer than about 25 seconds may be cut off.
  • Let two versions run at once. For a short time, old and new share your database. Renaming a column the old version still reads breaks it mid-update. Add first, switch over, remove later.
Background workers and cron jobs

Only web services have the gate. Background workers have no port and no health check, so a new worker version counts as ready as soon as its process is running. Cron jobs run on a schedule rather than continuously, so there's nothing to roll.