Problem Statement: CI/CD as a Distributed Control System
Frames the pipeline as four cooperating planes rather than a single build server, and names the correctness invariants that separate a delivery platform from a script.
Problem statement
Design a continuous integration and continuous delivery platform that triggers on every commit, runs builds and tests on an elastic fleet of ephemeral runners, produces immutable versioned artifacts, and promotes those artifacts through staging to production behind automated gates, approvals, canary analysis, and one-command rollback. The system must serve thousands of engineers committing concurrently, absorb bursty merge traffic, and never lose the audit trail of what shipped, why, and who authorized it.
The interview trap is drawing one box labeled "Jenkins". A production CI/CD platform is a distributed control system with four planes that fail differently and scale differently.
- Control plane: trigger resolution, pipeline-as-code compilation, scheduling, run/stage/job state machines, approvals, and the durable event log. It owns correctness.
- Execution plane: the ephemeral runner fleet, isolation (containers or microVMs), resource bin-packing, cache locality, and log streaming. It owns throughput and cost.
- Artifact plane: content-addressed build cache, OCI artifact registry, test reports, SBOMs, and signed provenance. It owns immutability and supply-chain trust.
- Delivery plane: environment model, promotion gates, deployment controllers, canary analysis, rollback, and GitOps reconciliation. It owns production safety.
Why this is distinctive
A build can be retried; a production deployment cannot be un-seen by users. The design therefore separates run progress (retryable, eventually advancing) from environment ownership (a fenced, single-writer invariant): only one deployment controller may mutate one environment at a time, and rollback is itself a forward deployment of a previously verified immutable artifact, never an in-place mutation.
Public evidence shows the category operates at extreme cadence. Amazon has publicly described deployment cadence on the order of one production deployment every eleven seconds across its engineering organization at AWS re:Invent, and Netflix open-sourced Spinnaker in 2015 after using it internally to automate thousands of daily microservice deployments with automated canary analysis. Those are cited public statements about those companies, not requirements for our design.
For capacity planning this answer explicitly assumes a mature platform with 6,000 engineers, 15,000 push events per day, 40,000 pipeline runs per day, 250,000 jobs per day, and 3,000 concurrently executing jobs at a 5x peak. Unless tied to a named public source, every number in this answer is a stated design assumption, target, or budget.
Key Highlights
- •CI/CD is four planes: control, execution, artifact, and delivery; each scales and fails differently.
- •Run progress is retryable; environment mutation is a fenced single-writer invariant.
- •Rollback is a forward deployment of a previously verified immutable artifact, never live mutation.
- •Amazon's publicly stated ~11.6-second deployment cadence and Netflix's 2015 Spinnaker open-sourcing prove the category's scale.
- •All uncited capacity numbers (6,000 engineers, 40,000 runs/day, 3,000 concurrent jobs) are explicit design assumptions.
Section Rescue Kit
Buzzwords to use:
Safe statements:
- "Let me separate the control plane that owns state from the execution plane that owns throughput before choosing any technology."
- "I will treat rollback as a forward deployment of an immutable artifact, which changes how I model environments."