Problem Statement: A Stream-Driven Personalization Engine
Frames the system as four planes: event ingestion, streaming feature computation, low-latency serving, and the closed learning loop.
Problem statement
Design a recommendation system that ingests user clickstream, watch, search, cart, and rating events continuously, updates session-level features and model state within seconds, and serves personalized top-K suggestions in under 100 ms p99. The system must tolerate partial and delayed events, re-rank candidates on the fly, and improve itself through a governed feedback loop.
This is not a batch nightly-recommendation job. A user who clicks a camping tent at 20:14 expects the next refresh to show sleeping bags, not yesterday's profile. Therefore the design separates four planes: (1) the ingestion plane that durably captures every interaction with ordering and replay; (2) the streaming compute plane that maintains session windows, real-time features, and online model updates; (3) the serving plane that retrieves candidates with approximate nearest neighbor search and ranks them under a hard latency budget; (4) the learning plane that closes the loop from impression to click to label to retraining.
Why stream-based is the distinctive requirement
A batch pipeline gives you stale personalization: features computed hours ago miss the current session intent. A pure online model with no durable stream gives you no replay, no audit, no consistent training/serving features. The interview-winning answer wires the stream as the spine: every feature, label, and model update is derived from the same replayable event log, so training/serving skew is bounded and failures are recoverable by replay.
Public operating baseline versus design assumptions
Public, company-reported figures establish that this category runs at extreme scale: Netflix has stated that a large majority of watched hours come from its recommendations, and Amazon has long been cited as attributing roughly a third of consumer revenue to recommendation-driven purchases. YouTube published the canonical two-stage deep network for candidate generation and ranking in 2016, and Meta published the Deep Learning Recommendation Model in 2019. These are cited context, not requirements for our fictional system.
For capacity planning this answer explicitly assumes: 300 million registered users, 40 million DAU, 8 million concurrent at peak, 2 billion recommendation requests per day, 50,000 interaction events per second average with a 5x peak, and a catalog of 100 million items. Unless tied to a citation, every number is a stated assumption, budget, or target.
Key Highlights
- •The stream is the spine: features, labels, and model updates all derive from one replayable event log.
- •Four planes: ingestion, streaming compute, low-latency serving, and the governed learning loop.
- •Session intent must influence the next refresh, which batch-only pipelines cannot deliver.
- •Assumed scale: 40M DAU, 8M concurrent, 2B rec requests/day, 50K events/s average with 5x peak.
- •Public figures (Netflix share of watch time, Amazon revenue attribution) are context, not requirements.
Section Rescue Kit
Buzzwords to use:
Safe statements:
- "I will separate what must react in seconds from what may refresh in hours, then place each on the right plane."
- "Before choosing technology, let me fix the event contract that every downstream consumer shares."