Design an E-Commerce Recommendation Engine

Hard45 min
1 / 30
understanding•11 min read

Problem Statement: Personalization as a Two-Plane ML Serving System

Frames the recommendation engine as an online serving plane coupled to an offline learning plane, not a single model box.

Problem statement

Design a recommendation engine for a large e-commerce platform that suggests relevant products across surfaces: the personalized home grid, product detail pages (similar items, frequently bought together), cart cross-sell, and re-engagement emails. The system must track user interactions (views, clicks, add-to-cart, purchases), compute item-to-item and user-to-item affinities, serve personalized lists with low latency, adapt to shifting user intent within a session, and prove effectiveness through controlled experimentation.

The defining structural fact is that a recommendation engine is two planes, not one service. The learning plane ingests billions of interactions, trains candidate-generation and ranking models, builds embedding indexes, and materializes features. The serving plane answers each page render in under 200 milliseconds: retrieve hundreds of candidates from a 100-million-item catalog, assemble features, score, re-rank under business constraints, and respond. These planes run on different timescales, different consistency rules, and different failure semantics, and the design must keep them coupled only through versioned artifacts and event streams.

Why the problem is distinctive

A CRUD backend answers a question about the current database state. A recommendation engine answers a question about predicted future behavior under a hard latency budget, using artifacts that are hours to days stale by construction. The answer therefore separates three concerns: correctness of the serving path (latency, availability, inventory truth), quality of the learning path (offline metrics, training data integrity, reproducibility), and business safety (never recommend out-of-stock items at the wrong price, honor merchandising rules, keep experiments statistically valid).

Public evidence shows the category operates at enormous scale. Amazon's 2003 IEEE Internet Computing paper described item-to-item collaborative filtering serving 29 million customers and 41 million items, with online computation scaling with user activity rather than catalog size; a widely cited McKinsey analysis attributes roughly 35 percent of Amazon's purchases to its recommendation features. Netflix states that about 80 percent of watched hours come from its recommendations and runs over a thousand concurrent experiments. Pinterest's PinSage paper (KDD 2018) reports graph convolutions over 3 billion nodes and 18 billion edges. These are published company figures and context, not targets for our fictional system.

For capacity planning, this answer explicitly assumes a mature platform with 50 million daily active users, a 100-million-SKU catalog, 2 billion recommendation requests per day, 2 billion interaction events per day, and a five-times event peak. Unless a number is tied to a citation, it is a stated design assumption, target, or budget.

The four architectural planes

  1. Collection plane: client beacons and server-side events capturing impressions, clicks, cart, purchase, and dwell with deduplication and sessionization.
  2. Learning plane: batch and streaming training, candidate-generation models, ranking models, embedding builds, and the experiment-driven release loop.
  3. Feature and index plane: offline feature materialization, online feature store, embedding index (ANN), and catalog/inventory projections.
  4. Serving plane: candidate retrieval, feature assembly, ranking, business re-ranking, response caching, and serving logs that close the feedback loop.

A strong interview answer keeps these planes separate. It allows the learning plane to stall without breaking the serving plane, and it lets the serving plane degrade to popularity-based results without corrupting training data.

Key Highlights

  • •The engine is two planes: a latency-critical serving plane and an eventually improving learning plane, coupled only through versioned artifacts and event streams.
  • •Serving must answer in under 200 ms over a 100M-item catalog; learning may take hours.
  • •Public anchors: Amazon item-to-item CF (2003, 29M customers, 41M items), ~35% of Amazon purchases attributed to recommendations (McKinsey), Netflix ~80% of watched hours from recommendations, Pinterest PinSage over 3B nodes and 18B edges.
  • •All uncited scale numbers in this answer are explicit assumptions: 50M DAU, 100M SKUs, 2B requests/day, 2B events/day, 5x peak.
  • •A degraded popular-items response is a successful serving outcome even when it is a weak personalization outcome.
Lead With the Two-Plane Split
State in the first two minutes that serving and learning are separate planes coupled only through versioned artifacts and streams. This instantly distinguishes an ML-systems architecture from a generic web backend.
Do Not Draw One Model Box
A design where a single model service both trains on clicks and answers page renders fails on latency, fails on freshness, and fails on blast radius. It will not survive a serious interview.

Section Rescue Kit

Buzzwords to use:

Candidate GenerationFeedback Loop

Safe statements:

  • "I will separate what must answer in 200 milliseconds from what may take hours to improve."
  • "Before choosing models, let me define which plane owns each decision and how artifacts move between them."
Design an E-Commerce Recommendation Engine - System Design | WinJob | WinJob