Key idea
A migration is a small, numbered file that changes the schema by one step. Migrations live in git next to the code, run in order, and are recorded in the database, so every environment reaches the same schema by running the same files.
Why not just change the database?
You could open a SQL console on production and add the column by hand. Then your laptop, staging and production each end up slightly different, and nobody can say which changes production has.
The bill arrives on release day: code that expects due_date ships to a database without it, and every request that touches tasks fails.
How migrations work
- Numbered files.
001_create_tasks, then002_add_due_date. Order matters: 002 changes a table that 001 creates. - Up and down. Each migration has an up, which makes the change, and a down, which undoes it.
- A record in the database. A table, often called
schema_migrations, lists what has run. The tool compares it with the files and runs only what's missing, so running it twice changes nothing. - One step at a time. Where the database allows it, each migration runs in a transaction, a group of changes that all happen or none do, so a failed one leaves nothing half-done.
- The same files everywhere. Laptop, staging and production run the same migrations in the same order. A migration is reviewed in the same pull request (a proposed change others review before it's merged, Module 7) as the code that needs it.
When they run
Run migrations once per release, before the new code needs them, from one place: a release step (a command your deploy runs once before starting the new version) or a one-off job. Running them from every copy of the app as it starts means several copies race to change the same table.
Down isn't a backup
A down that drops a column deletes the data in it, and running up again won't bring the data back. Take a backup before any migration that removes something. Treat down as a tool for development, and for undoing a step that has only just gone out.
Changing a schema with no downtime
During a release, the old and new versions of your app briefly run side by side against the same database, so every schema change must work for both. The pattern is expand, then contract: add the new column first, move the code over, and remove the old column in a later release.
Expand and contract, step by step
Say you want to rename title to name. A single rename breaks the old version, which still reads title while the new one starts. Instead:
- Expand. A migration adds
name. The code writes both columns and still readstitle. - Backfill. Copy
titleintonamefor existing rows, in batches if the table is large. - Switch. The code reads and writes only
name. - Contract. In a later release, once no running version reads
title, a migration drops it.
Each release works with the schema before and after its own migration. It takes more releases, and in return no request fails along the way. Lesson 5.7.2 shows the same rule for rolling updates on ComputeSphere.
Check yourself