Architecture Reference // 2026

Progressive
Delivery

Deploy to infrastructure. Release to users. Control the distance between the two.
8 domains  ·  27 rules.

BLD
→
CAN
→
VAL
→
PRD

BUILD  ·  CANARY  ·  VALIDATE  ·  PROMOTE

01 // Foundation

Progressive delivery is a risk management system

A deploy to 100% of users simultaneously is a binary bet — it works, or it fails for everyone at once. Progressive delivery turns that binary bet into a controlled sequence with an exit at every stage. The goal is not slower shipping — it is smaller blast radius per unit of change.

Deploy and release are different events
Code is deployed to infrastructure — servers, containers, pods. That is an engineering event. Code is released to users. That is a product event. Progressive delivery is the layer that manages the distance between these two. Collapsing them into a single operation removes all control over blast radius.
Not every change needs a canary — but you need a framework for which ones do
A typo fix doesn't need staged rollout. A new payment flow does. The discipline is a consistent risk classification at the point of merge: what is the scope of breakage if this is wrong? High-risk changes get progressive delivery automatically. Low-risk changes don't accumulate ceremony they don't need.
Progressive delivery requires investment in observability first
A canary with no metrics to evaluate is just a slow deploy. The ability to promote or rollback based on real signal depends entirely on instrumentation that exists before the release begins. You cannot build progressive delivery on top of poor observability — the promotion criteria have nothing to read from.
02 // Canary Releases

Start at 1%. Not 10%.

A canary release routes a small slice of production traffic to the new version while the stable version handles the rest. It catches real-world edge cases that staging never will — at a blast radius small enough to absorb the failure. The key discipline is defining promotion criteria before the canary starts, not while you're watching it.

01
1% first — resist the urge to start at 10%
The first canary slice should be small enough that a catastrophic failure affects almost nobody. 1% of production traffic finds real-world conditions staging never replicates. Teams who start at 10% "to be faster" are trading speed for the blast-radius reduction that makes canaries worth having.
02
Define promotion criteria before the canary starts
The metrics, thresholds, and soak time that trigger promotion are written down before data collection begins. If error rate stays below X and p99 latency stays below Y for Z minutes, promote. Anything else is a rollback. Defining these after seeing data is peeking — it produces the decision you wanted, not the decision that is correct.
03
Canary traffic must be representative — not cherry-picked
Random slice of production traffic — not internal users only, not a specific region, not low-traffic hours. A canary on non-representative traffic passes and then fails at 100%. The whole point is catching real-world edge cases: high-volume accounts, unusual request shapes, the edge of your input space.
04
Automatic rollback must be faster than human reaction
A canary that requires a human to notice a degraded metric, deliberate, and act is not a safety system — it is a delayed big-bang deploy with extra steps. Unambiguous degradation triggers rollback automatically. Human judgment is for ambiguous signals. Clear threshold violations should never wait for anyone.
03 // Blue / Green

Instant rollback — not gradual exposure

Blue/green maintains two identical production environments. Traffic flips 100% between them in seconds. This is not a canary substitute — it is a fast-switch mechanism. The trade-off: no blast-radius reduction on the way in, but instant, tested escape if something goes wrong after cutover.

Blue/green gives you instant rollback — not progressive exposure
The value of blue/green is not in how you get to 100% — it is in how fast you get back to 0%. A rollback is a traffic pointer change, not a redeploy. This makes it the correct mechanism for the final cutover step, not the validation step. Canary validates; blue/green escapes.
Two full environments is the real cost
For stateless services, running two identical environments is manageable infrastructure overhead. For stateful services — databases, persistent queues, shared caches — the complexity rises significantly. The new environment must read from the same data store as the old. Schema changes must be backward compatible with both simultaneously.
Canary validates. Blue/green escapes. Use both.
The patterns are complementary. Run a canary to validate the change progressively against real traffic. Once confidence is established, use a blue/green flip for the final 0% → 100% cutover — and keep the blue environment available as an instant rollback for the soak period. Choosing one over the other is a false trade-off.
04 // Traffic Management

Traffic shifting happens at a specific layer. Know which one.

The mechanism that routes traffic between versions determines the granularity, latency, and operational complexity of your progressive delivery. Choosing the wrong layer creates invisible constraints — you may discover mid-incident that you cannot shift traffic as fast as you assumed.

Each layer has different capabilities and costs
Load balancer: coarse weight-based splitting, sub-second changes, no per-request context. Service mesh (Istio, Linkerd): header-based routing, fine-grained per-path rules, operational overhead. CDN/edge: geographic splitting, global propagation delay. Feature flags: per-user, per-tenant, application-level — the most granular but requires code instrumentation. Each is correct for different requirements.
Header-based routing enables deterministic pre-production testing
A traffic management layer that routes on a request header — X-Canary: true — lets QA environments, internal users, and specific tenants be explicitly routed to the new version regardless of the traffic percentage. This is how you validate before the percentage rollout begins. Without it, your only option is waiting for the random slice to hit the right conditions.
Sticky routing during a canary must be intentional
A user who hits the canary on request 1 and the stable version on request 2 sees inconsistent behavior — and your metrics measure a mix of both experiences. Sticky routing by user_id hash ensures each user stays on one version for the duration of the canary. But sticky routing also means canary exposure compounds over time rather than being random per-request. Know which behavior you want before the canary starts.
Traffic shifting layers — coarse to fine
Load balancer — weight-based
DNS — weighted routing
Service mesh — header / path rules
Feature flags — per user / tenant
CDN edge — geographic
Kubernetes — replica weight
05 // Rollback Strategy

Rollback is not a failure state. It is the system working as designed.

Teams that treat rollback as an admission of failure will delay it, avoid it, and attempt to fix forward instead. Fix-forward under pressure, without understood root cause, is how SEV2s become SEV1s. Rollback is a first-class, frequently practiced operation — not a last resort.

Roll back on unknown-cause degradation. Always.

Roll forward when: the fix is simple, low-risk, and faster than a rollback. Roll back when: the root cause is unclear, customer impact is ongoing, or the fix requires investigation. For any ambiguous degradation during a canary, the default is rollback — not investigation under fire. You can investigate cleanly once users are unblocked.

Test your rollback path as often as your deploy path

A rollback procedure that has never been exercised in production will fail the first time it is needed — under pressure, at the worst moment. Include rollback in regular deployment practice. A well-practiced rollback takes seconds. An untested one takes minutes. During a SEV1, minutes are not recoverable.

Rollback criteria are written before the deploy begins

The same discipline that applies to canary promotion applies to rollback: define the trigger conditions before the deploy. "Error rate exceeds 2% for 5 consecutive minutes" is a rollback trigger. "Something feels off" is not. Pre-defined criteria remove the decision from the moment of pressure — when cognitive load is highest and bias toward optimism is strongest.

Database rollbacks are different — plan separately

Application rollback is cheap. Schema rollback is often impossible — a dropped column or renamed table cannot be un-migrated without data loss. The expand/contract pattern exists specifically because schema changes must be written to survive a rollback of the application layer. Never assume a schema migration is reversible.

06 // SLO-Gated Promotion

The deployment pipeline reads from your observability stack

Promotion decisions made by humans looking at dashboards are susceptible to optimism bias, time pressure, and incomplete information. At any meaningful deploy frequency, human-in-the-loop promotion does not scale. The deployment pipeline should read SLO state as a first-class input to every promotion decision.

1

Manual promotion with metric dashboards

An engineer reviews error rate and latency before promoting. Better than nothing. Susceptible to cognitive bias — engineers look for reasons to promote, not reasons to pause. Doesn't scale past a few deploys per day. Promotions that happen at 11pm are not checked the same as promotions that happen at 11am.

2

Automated checks as promotion prerequisites

Promotion requires all checks to pass: error rate below threshold, p99 latency within SLO, no active alerts on the canary, SRM check clean. Humans can override but cannot skip. This is the correct default for any team shipping more than a handful of times per day. The pipeline does the work; humans review exceptions.

3

Error budget as the deployment gate

Promotion is gated on SLO error budget remaining — not just current snapshot metrics. A service burning through its error budget at 3x the normal rate does not get promoted to 100%, even if the current metric window looks fine. The pipeline integrates directly with your SLO tooling. Deploys slow down as reliability degrades — automatically.

07 // Database & Migrations

The database is where progressive delivery breaks

Application code can be progressively shifted between versions. Database schemas cannot. A migration applied while two application versions are running simultaneously must be backward compatible with both. Most progressive delivery failures are not in the application layer — they are in assumptions about the database that were never made explicit.

Expand/contract is the only safe schema migration pattern
Expand: add the new column, table, or index while the old schema still works. Deploy the new application version. Both old and new application code run against the expanded schema simultaneously. Contract: once the old version is fully retired and no traffic touches the old path, remove the old column in a separate deployment. The contraction step is always a separate release — never bundled with the expand.
Data migrations and schema migrations are separate releases
Migrating data — backfilling a new column, rewriting rows, re-encoding values — is a different operation from changing the schema. Data migrations run in the background with no immediate traffic impact. Schema migrations affect query plans, index builds, replication lag, and table locks immediately. Bundling them into the same deployment turns a controlled migration into an uncontrolled risk event.
Feature-flag the new read path before migrating
Before reading from a new schema in production, gate the read path behind a feature flag. Deploy the new code with the flag off. Validate the schema change against live data without serving it. Enable the flag progressively — 1%, 10%, 100%. This gives you a rollback path on the read without a schema rollback — which is often impossible and always expensive.
Test migrations against a production-scale dataset
A migration that takes 3 seconds on a staging database with 10,000 rows takes 45 minutes on production with 50 million rows — and holds a table lock the entire time. Migration duration and lock behavior must be tested against a representative data volume before the production window. Pg shadow clones, MySQL staging replicas, and cloud snapshot restores exist for exactly this reason.
08 // DORA & Culture

High deploy frequency and low failure rate move together

The intuition that deploying more often means more breakage is wrong in high-performing teams. Small, frequent, progressive deployments have lower blast radius, faster detection, faster rollback, and a cleaner causal signal than large, infrequent ones. The DORA research makes this unambiguous: the risk per deploy decreases as deploy frequency increases.

DORA metrics are the scoreboard

Four metrics measure the outcome of your delivery system: Deployment Frequency (how often), Lead Time for Changes (how fast from commit to production), Change Failure Rate (what fraction of deploys cause incidents), Time to Restore Service (how quickly you recover). High performers deploy multiple times per day with under 5% change failure rate and restore in under an hour. Progressive delivery is the primary enabler.

Progressive delivery is a product capability, not just an ops one

The ability to ship to 1% of users, measure the result, and decide whether to proceed is a product capability as much as a deployment one. Teams with this infrastructure can validate product hypotheses in hours, run experiments on production traffic, and respond to real-world signal before a full rollout commits the team to a direction. It is the foundation of a data-driven product process.

Invest in the deployment pipeline like production infrastructure

Flaky tests, slow builds, manual promotion steps, and undocumented rollback procedures are technical debt that compounds with every deploy. The deployment pipeline is on the critical path of every feature that ships — treat it with the same engineering rigor as production services: SLOs, monitoring, runbooks, on-call. A 20-minute build is an invisible tax on every engineer, every day.

Feature freeze is a symptom, not a strategy

Teams that institute feature freezes before releases are compensating for the inability to deploy safely. A team with progressive delivery, automated rollback, and SLO-gated promotion does not need feature freezes — every deploy is already contained. If your team reaches for a freeze, the right question is: what would need to be true about our delivery system to make this unnecessary?

The mental model

Deploy to infrastructure. Release to users. They are different events.
A canary is not a slow deploy — it is a controlled experiment with an exit.
Define rollback criteria before you start. Not while you're deciding.
The database is where progressive delivery breaks. Use expand/contract.
High deploy frequency and low failure rate are not in tension. They move together.

Daniel Brasileiro