Zero downtime deployments: rolling, blue-green and canary compared
Zero downtime is not a deployment tool setting. It is a property of an application that can run two versions at once.
Safe migrations follow the expand/contract pattern: add the new structure, deploy code that writes to both old and new, backfill in batches, switch reads to the new structure, and only then remove the old one. Each step is independently deployable and reversible, so no release requires the schema and the code to change at the same instant.
| Operation | Lock impact (modern PostgreSQL) |
|---|---|
| ADD COLUMN (nullable, no default) | Safe — metadata only |
| ADD COLUMN with a constant default | Safe in PostgreSQL 11+ |
| ADD COLUMN with a volatile default | Rewrites the table — dangerous |
| CREATE INDEX | Blocks writes — use CONCURRENTLY |
| ADD CONSTRAINT ... NOT VALID then VALIDATE | Safe in two steps |
| ALTER COLUMN TYPE | Usually rewrites the table |
| DROP COLUMN | Fast, but breaks any deployed code still selecting it |
Split a breaking schema change into a sequence of individually safe, backwards-compatible deploys.
Renaming users.name -> users.full_name
1. Expand ALTER TABLE users ADD COLUMN full_name text;
2. Dual write Deploy code writing both name and full_name
3. Backfill UPDATE in batches until full_name is complete
4. Switch Deploy code reading full_name (still writing both)
5. Contract Stop writing name; ALTER TABLE users DROP COLUMN name;Five deploys instead of one. Every step is reversible, no step requires the database and the application to be in lockstep, and a rollback at any point leaves a working system. That is what makes it worth the extra ceremony on anything user-facing.
-- Batched backfill: bounded work per statement, pauses between batches
DO $$
DECLARE rows_updated integer;
BEGIN
LOOP
UPDATE users SET full_name = name
WHERE id IN (SELECT id FROM users WHERE full_name IS NULL LIMIT 5000);
GET DIAGNOSTICS rows_updated = ROW_COUNT;
EXIT WHEN rows_updated = 0;
COMMIT;
PERFORM pg_sleep(0.1);
END LOOP;
END $$;Additive, safe migrations, yes — as a separate step before the application rolls out. Destructive or long-running ones should be run deliberately with someone watching.
For additive changes, deploy the previous code — the extra column is harmless. For destructive changes there is no rollback without a restore, which is precisely why expand/contract defers destruction to the very last step.
They generate the SQL; they do not understand your traffic or lock behaviour. Always read the generated SQL before applying it to production.
Restore a recent production backup to a staging database of similar size and run it there with representative concurrent load. Row count alone is not enough — locks only bite under traffic.
ROVQIX Engineering
Engineering team, ROVQIX
The ROVQIX engineering team builds and maintains web platforms, APIs and infrastructure for clients across SaaS, ecommerce and enterprise. These notes come out of real production work — deploys, incidents, migrations and audits.
ROVQIXdesigns and builds production web platforms — Next.js front ends, Node.js APIs and the infrastructure behind them. Tell us what you're building and we'll scope it with you.
Zero downtime is not a deployment tool setting. It is a property of an application that can run two versions at once.
You do not have backups. You have restores — and you only know whether you have those if you have done one this quarter.
Adding an index is easy. Knowing which one, in which column order, and which existing indexes to delete is the part that changes query times.
No spam. Just the occasional case study and craft breakdown.