Deploy to infrastructure. Release to users. Control the distance between the two.
8 domains · 27 rules.
BUILD · CANARY · VALIDATE · PROMOTE
A deploy to 100% of users simultaneously is a binary bet — it works, or it fails for everyone at once. Progressive delivery turns that binary bet into a controlled sequence with an exit at every stage. The goal is not slower shipping — it is smaller blast radius per unit of change.
A canary release routes a small slice of production traffic to the new version while the stable version handles the rest. It catches real-world edge cases that staging never will — at a blast radius small enough to absorb the failure. The key discipline is defining promotion criteria before the canary starts, not while you're watching it.
Blue/green maintains two identical production environments. Traffic flips 100% between them in seconds. This is not a canary substitute — it is a fast-switch mechanism. The trade-off: no blast-radius reduction on the way in, but instant, tested escape if something goes wrong after cutover.
The mechanism that routes traffic between versions determines the granularity, latency, and operational complexity of your progressive delivery. Choosing the wrong layer creates invisible constraints — you may discover mid-incident that you cannot shift traffic as fast as you assumed.
X-Canary: true — lets QA environments, internal users, and specific tenants be explicitly routed to the new version regardless of the traffic percentage. This is how you validate before the percentage rollout begins. Without it, your only option is waiting for the random slice to hit the right conditions.user_id hash ensures each user stays on one version for the duration of the canary. But sticky routing also means canary exposure compounds over time rather than being random per-request. Know which behavior you want before the canary starts.Teams that treat rollback as an admission of failure will delay it, avoid it, and attempt to fix forward instead. Fix-forward under pressure, without understood root cause, is how SEV2s become SEV1s. Rollback is a first-class, frequently practiced operation — not a last resort.
Roll forward when: the fix is simple, low-risk, and faster than a rollback. Roll back when: the root cause is unclear, customer impact is ongoing, or the fix requires investigation. For any ambiguous degradation during a canary, the default is rollback — not investigation under fire. You can investigate cleanly once users are unblocked.
A rollback procedure that has never been exercised in production will fail the first time it is needed — under pressure, at the worst moment. Include rollback in regular deployment practice. A well-practiced rollback takes seconds. An untested one takes minutes. During a SEV1, minutes are not recoverable.
The same discipline that applies to canary promotion applies to rollback: define the trigger conditions before the deploy. "Error rate exceeds 2% for 5 consecutive minutes" is a rollback trigger. "Something feels off" is not. Pre-defined criteria remove the decision from the moment of pressure — when cognitive load is highest and bias toward optimism is strongest.
Application rollback is cheap. Schema rollback is often impossible — a dropped column or renamed table cannot be un-migrated without data loss. The expand/contract pattern exists specifically because schema changes must be written to survive a rollback of the application layer. Never assume a schema migration is reversible.
Promotion decisions made by humans looking at dashboards are susceptible to optimism bias, time pressure, and incomplete information. At any meaningful deploy frequency, human-in-the-loop promotion does not scale. The deployment pipeline should read SLO state as a first-class input to every promotion decision.
An engineer reviews error rate and latency before promoting. Better than nothing. Susceptible to cognitive bias — engineers look for reasons to promote, not reasons to pause. Doesn't scale past a few deploys per day. Promotions that happen at 11pm are not checked the same as promotions that happen at 11am.
Promotion requires all checks to pass: error rate below threshold, p99 latency within SLO, no active alerts on the canary, SRM check clean. Humans can override but cannot skip. This is the correct default for any team shipping more than a handful of times per day. The pipeline does the work; humans review exceptions.
Promotion is gated on SLO error budget remaining — not just current snapshot metrics. A service burning through its error budget at 3x the normal rate does not get promoted to 100%, even if the current metric window looks fine. The pipeline integrates directly with your SLO tooling. Deploys slow down as reliability degrades — automatically.
Application code can be progressively shifted between versions. Database schemas cannot. A migration applied while two application versions are running simultaneously must be backward compatible with both. Most progressive delivery failures are not in the application layer — they are in assumptions about the database that were never made explicit.
The intuition that deploying more often means more breakage is wrong in high-performing teams. Small, frequent, progressive deployments have lower blast radius, faster detection, faster rollback, and a cleaner causal signal than large, infrequent ones. The DORA research makes this unambiguous: the risk per deploy decreases as deploy frequency increases.
Four metrics measure the outcome of your delivery system: Deployment Frequency (how often), Lead Time for Changes (how fast from commit to production), Change Failure Rate (what fraction of deploys cause incidents), Time to Restore Service (how quickly you recover). High performers deploy multiple times per day with under 5% change failure rate and restore in under an hour. Progressive delivery is the primary enabler.
The ability to ship to 1% of users, measure the result, and decide whether to proceed is a product capability as much as a deployment one. Teams with this infrastructure can validate product hypotheses in hours, run experiments on production traffic, and respond to real-world signal before a full rollout commits the team to a direction. It is the foundation of a data-driven product process.
Flaky tests, slow builds, manual promotion steps, and undocumented rollback procedures are technical debt that compounds with every deploy. The deployment pipeline is on the critical path of every feature that ships — treat it with the same engineering rigor as production services: SLOs, monitoring, runbooks, on-call. A 20-minute build is an invisible tax on every engineer, every day.
Teams that institute feature freezes before releases are compensating for the inability to deploy safely. A team with progressive delivery, automated rollback, and SLO-gated promotion does not need feature freezes — every deploy is already contained. If your team reaches for a freeze, the right question is: what would need to be true about our delivery system to make this unnecessary?
Deploy to infrastructure. Release to users. They are different events.
A canary is not a slow deploy — it is a controlled experiment with an exit.
Define rollback criteria before you start. Not while you're deciding.
The database is where progressive delivery breaks. Use expand/contract.
High deploy frequency and low failure rate are not in tension. They move together.