Problem Statement: A Recommender That Learns From Every Impression
Frames the system as a closed-loop decision engine, not a batch scoring pipeline, and isolates concurrency as the differentiating hard problem.
Problem statement
Design a recommender that treats every user impression as a bandit trial: for each request it observes user and context features, selects items to show from a candidate pool, emits the decision with a logged propensity, observes immediate reward (click, dwell, conversion, or none), and folds that reward back into the policy so future decisions improve. The system must support epsilon-greedy and Thompson sampling baselines, contextual bandits (LinUCB-style) for context-aware decisions, and leave a credible path toward session-level deep RL.
This is not a batch ranking job with a nightly model. A bandit recommender is a closed-loop control system: the serving path, the logging path, and the learning path form a cycle that runs continuously, and every component of that cycle is on the request-time critical path of the product. Netflix publicly describes its artwork personalization as a contextual-bandit-style system that collects interaction data, learns preferences, and deploys updated models continuously; YouTube's research team has published on production REINFORCE recommenders with top-K off-policy correction, showing the industry direction beyond bandits.
Why the problem is distinctive
A classical recommender answers "what does this user probably like?" using historical data. A bandit recommender additionally answers "what should we show to learn?" It must deliberately take actions whose immediate reward may be lower in order to reduce uncertainty, while bounding the revenue cost of that exploration. Three properties make the design hard:
First, feedback is biased by the system's own actions. The only rewards observed are for items the policy chose to show, so naive retraining amplifies selection bias. Every impression must log the probability (propensity) with which the action was taken, enabling off-policy evaluation without a full A/B test.
Second, rewards are delayed and noisy. A click arrives seconds later, a dwell-time signal minutes later, a subscription days later. The reward-join pipeline must attribute late events to impressions with watermarks and tolerances, and the update loop must tolerate label noise without oscillating.
Third, and this is where most shallow answers stop: concurrency at massive user scale. Millions of concurrent users simultaneously generate rewards that update shared arm statistics and model parameters, while millions of read requests need low-latency access to those same parameters. The parameter plane is a hot, multi-reader, high-write-rate state store with convergence requirements, not just a model file on disk. This answer details that bridge explicitly.
Public operating baseline versus design assumptions
Netflix reported approximately 282.7 million paid memberships in its Q4 2024 results and 184 billion hours viewed in the second half of 2023 in its engagement report. Spotify reported 675 million monthly active users for Q4 2024. LinkedIn has publicly stated more than 1 billion registered members. These are company-reported figures used as context for scale. For capacity planning, this answer assumes a fictional mature product: 100 million DAU, 2.4 billion recommendation requests per day, 48 billion impressions per day, and a 5x event peak. Unless tied to a named company publication, every number here is a stated design assumption.
The four architectural planes
- Decision plane: candidate retrieval, feature assembly, bandit scoring, exploration sampling, filters, and response assembly under a 50 ms p99 budget.
- Feedback plane: impression and reward event logging with propensity, policy version, and experiment metadata; delayed reward joins.
- Learning plane: per-shard parameter aggregation, hourly batch training, off-policy evaluation, and gated promotion.
- Governance plane: experiment registry, exploration budgets, content guardrails, fairness audits, privacy controls, and rollback.
A strong interview answer keeps these planes separate: the decision plane degrades without the learning plane, the learning plane never writes directly into the serving hot path, and governance constraints are enforced at decision time, not audited afterwards.
Key Highlights
- •The system is a closed loop: decide, log with propensity, join delayed rewards, update parameters, redeploy.
- •Selection bias is the core statistical hazard; every decision logs the probability it was taken with.
- •Concurrency on shared parameters at 100M DAU is the differentiating engineering problem, not the algorithm choice.
- •Company figures from Netflix, Spotify, and LinkedIn set context; all scale targets here are explicit assumptions.
- •Four planes: decision, feedback, learning, governance. Each degrades independently.
Section Rescue Kit
Buzzwords to use:
Safe statements:
- "I will separate the decision path from the learning path: serving must stay fast even when training is broken."
- "Before choosing an algorithm, let me define what gets logged, because the log schema determines what learning is possible."