Running database migrations without taking the site down
The migration that takes your site down is almost never the complicated one. It is the ALTER TABLE that took a lock nobody expected.
Rolling updates replace instances gradually and are the default on most platforms. Blue-green runs two full environments and switches traffic atomically, giving instant rollback. Canary sends a small percentage of traffic to the new version first, catching problems with minimal exposure. All three require backwards-compatible schemas and graceful shutdown to actually be zero downtime.
| Rolling | Blue-green | Canary | |
|---|---|---|---|
| Extra infrastructure | Minimal | 2× during switch | Small |
| Rollback speed | Another rolling update | Instant traffic switch | Shift traffic back |
| Blast radius on failure | Growing during rollout | All users after switch | Small percentage |
| Complexity | Low | Medium | High |
| Needs metrics automation | No | No | Yes |
Most teams should run rolling updates well before attempting canary. A canary without automated metric analysis and an abort rule is a rolling update with extra steps.
It must tolerate the old and new versions running at the same time against the same data.
They separate deploying code from releasing behaviour. You ship the new path disabled, enable it for internal users, then a percentage, then everyone — and disable it instantly if something breaks, without a deploy at all.
Define the abort condition in advance and write it down. 'Error rate above 1% for two minutes, or p99 latency 50% worse than baseline' is actionable at 2am. 'It looks a bit off' is not.
Not usually. Blue-green earns its keep when rollback speed is critical and you can afford double capacity briefly. Rolling updates with good probes cover most web applications.
Long enough to gather statistically meaningful data on your traffic volume — typically 15 minutes to a few hours. High-traffic services can be confident far faster than low-traffic ones.
Partially — a process manager can start the new version, wait for readiness and switch traffic locally. You still have a window during restarts, and no protection against the machine itself failing.
Deploy the migration separately and before the code that depends on it, using expand/contract so both versions work against the intermediate schema.
ROVQIX Engineering
Engineering team, ROVQIX
The ROVQIX engineering team builds and maintains web platforms, APIs and infrastructure for clients across SaaS, ecommerce and enterprise. These notes come out of real production work — deploys, incidents, migrations and audits.
ROVQIXdesigns and builds production web platforms — Next.js front ends, Node.js APIs and the infrastructure behind them. Tell us what you're building and we'll scope it with you.
The migration that takes your site down is almost never the complicated one. It is the ALTER TABLE that took a lock nobody expected.
A pipeline that takes 25 minutes and fails randomly does not improve quality. It teaches the team to merge on red.
During an incident, the instinct is to find the cause. The job is to restore service. Those are different activities and the order matters.
No spam. Just the occasional case study and craft breakdown.