Design a ML Model Deployment & A/B Testing Platform

Hard45 min
1 / 30
understanding•11 min read

Problem Statement: A Production ML Deployment and Experimentation Control Plane

Frames the platform as a traffic-management and statistical-inference system, not merely a model-hosting service.

Problem statement

Design a platform that takes newly trained ML models, packages them as versioned container images or serialized artifacts, deploys them behind a traffic-splitting router, routes a configurable fraction of live inference requests to each model variant, ingests real-time business and latency metrics per variant, evaluates statistical significance or bandit reward signals, and either promotes the winning variant to full traffic or rolls back to the previous champion. The platform must store experiment logs for offline analysis, track per-model compute cost, and expose a UI for data scientists and product managers who are not engineers.

This is not a model-training problem. Training happens upstream in notebooks, pipelines, or managed services. This platform owns the moment a trained artifact crosses the boundary into production traffic. That boundary is where most ML value is destroyed: a model that improves offline AUC by two percentage points can degrade a business KPI by five percent if the training distribution has shifted, if the serving latency budget is violated, or if the traffic slice is too small to reach statistical significance before a product decision is forced.

Why the problem is distinctive

A traditional microservice canary deployment compares error rates and latency. An ML model deployment compares business outcomes that are noisy, delayed, and confounded. A recommendation model's click-through rate depends on the user's mood, the time of day, the inventory available, and the behavior of the other nine experiments running concurrently on the same page. The platform must therefore provide: deterministic traffic assignment that is stable per user across sessions; orthogonal experiment layers so that concurrent tests do not contaminate each other; a metrics pipeline that joins inference logs with downstream conversion events that may arrive hours or days later; and a statistical engine that supports both fixed-horizon hypothesis testing and sequential or bandit-based adaptive allocation.

The problem requires model container packaging and version registry, weighted traffic splitting for A/B or multi-armed bandit, real-time metrics ingestion covering latency and error rate and user-level business metrics, automated scale-up or rollback logic, safe fallback if a new model fails, scalability for many models and high QPS, low routing overhead, and a UI for non-engineers. Every one of these is designed explicitly below.

Public operating baseline versus design assumptions

Public evidence establishes that the category is operationally mature. Uber's Michelangelo platform, described in a 2017 engineering blog post, serves billions of predictions per day across dozens of models with a unified deploy-and-monitor pipeline. Netflix published details of its Metaflow framework and its experimentation infrastructure supporting thousands of concurrent A/B tests across 260 million subscribers. Airbnb's Bighead ML platform and its reporting pipeline handle experiment readouts across millions of nightly bookings. LinkedIn's XLNT experimentation platform runs tens of thousands of concurrent experiments on a member base exceeding one billion. These are cited company figures that establish the problem is real and solved at scale. They are not the requirements for our fictional system.

For capacity planning, this answer explicitly assumes a mature organization with 500 ML models registered, 120 models actively serving production traffic, 40 concurrent experiments at any time, 8,000 inference requests per second at global peak, and a five-times event multiplier during product launches or seasonal peaks. Unless a number is tied to a citation, it is a stated design assumption, target, budget, or illustrative threshold.

The four architectural planes

  1. Registry and packaging plane: model artifact storage, container image building, version metadata, lineage tracking, and approval gates.
  2. Traffic and deployment plane: the inference router, canary controller, experiment allocator, and rollback automation.
  3. Metrics and decision plane: real-time metric ingestion, statistical testing engine, bandit policy engine, promotion and rollback triggers.
  4. Governance and observability plane: experiment UI, cost tracking, audit logs, compliance, and alerting.

A strong interview answer keeps these planes separate. The registry plane must never block serving traffic. The metrics plane must never gate the inference hot path. The decision plane must operate on aggregated evidence, not raw event streams. And the governance plane must not require engineering access for routine experiment lifecycle operations.

Key Highlights

  • •The platform owns the boundary between a trained artifact and live production traffic, not the training itself.
  • •ML canary evaluation requires business-outcome comparison, not just error-rate and latency comparison.
  • •Public figures from Uber Michelangelo, Netflix Metaflow, and Airbnb Bighead establish the problem at real scale.
  • •The assumed mature organization runs 500 registered models, 120 serving, 40 concurrent experiments, 8K QPS peak.
  • •The architecture separates registry, traffic, metrics-decision, and governance into four independent planes.
  • •The inference hot path must never be blocked by the metrics or decision planes.
Lead With the Serving Boundary
State in the first two minutes that this platform owns the transition from trained artifact to live traffic. That instantly separates your answer from a model-training pipeline design and shows you understand where ML value is actually created or destroyed.
Do Not Design a Model Training Pipeline
A design that spends most of its time on GPU clusters, hyperparameter tuning, and training data pipelines has misunderstood the question. The problem asks for deployment, traffic splitting, metrics, and rollback. Training is upstream.

Section Rescue Kit

Buzzwords to use:

Champion-ChallengerOrthogonal Experiment Layers

Safe statements:

  • "I will separate model packaging from model serving from experiment evaluation, because each has different latency, consistency, and failure requirements."
  • "Before choosing infrastructure, let me define what a successful experiment looks like and what triggers an automatic rollback."
Design a ML Model Deployment & A/B Testing Platform - System Design | WinJob | WinJob