Learning paths / Deploy on ComputeSphere / Data: databases and storage

Stateless services and where state lives

Reading · 5 min · Module 5, lesson 1 of 546 min left in this module

Module 5 · Data: databases and storageLesson 1 of 5

Goal: Decide where each piece of your app's data belongs, and keep it out of the spherelet's own filesystem.

Key idea

Anything your app writes to its own filesystem disappears when the spherelet is replaced. Keep data somewhere built to keep it: a database, object storage, or a SphereStor volume.

The spherelet's filesystem is scratch space

Each spherelet starts with a fresh copy of the filesystem from your image. Your app can write to it, but those writes belong to that one spherelet and vanish when it's replaced.

That happens more often than you'd expect: on every restart and every redeploy that applies a change, after a crash, and when you scale. A new spherelet never sees files an old one wrote.

So a service that saves uploads to /app/uploads works in testing, then loses every file on the next deploy. With two spherelets it's worse: an upload lands on one, and the next request may reach the other.

Stateless is the goal

A stateless service keeps nothing it needs on its own disk. Any spherelet can answer any request, so you can replace, restart or add spherelets without losing anything. Most web services and workers can be written this way.

Temporary files are still fine: a file you write, use and delete in one request, or a cache you can rebuild. The rule is only that losing the file must never lose data.

Where state belongs

Kind of dataWhere it belongsExample
Records you query: users, ordersA databasePostgreSQL or MySQL from a hosted provider
Files and media, many or largeObject storageA bucket from your cloud or storage provider
Files one service keeps on diskA volume (SphereStor)A SQLite file, a search index
Values that are cheap to loseMemory or a cacheRendered pages, rate-limit counters

Object storage is a service that stores whole files by name over HTTP; "S3-compatible" means it speaks the same API as Amazon S3, so most libraries work with it.

ComputeSphere doesn't host databases or object storage. You bring them from a provider and connect with secrets (lesson 5.5.3). What it does provide is SphereStor: a volume that stays in your environment when the spherelet using it is replaced.

A volume or a database?

A volume looks like a folder, which makes it tempting for everything. But it attaches to one spherelet at a time, so it doesn't fit a service you want to scale out. Backups and recovery are also up to you.

A database handles connections from many spherelets at once and gives you queries and transactions (several changes that succeed or fail together). If more than one copy of your service, or another service, needs the data, use a database or object storage.

A quick test for any service

Imagine it running on three spherelets, then restarted. Would any request fail, or any data be missing? If yes, find the file or in-memory value that caused it and move it to one of the places in the table.

Variables and secrets are configuration, not data. Keep them in the service's settings rather than baked into the image, so the same image runs in every environment.

Check yourself

Your web service runs on 3 spherelets and needs to keep users' profile photos. Where should the photos go?

In the docs