Problem Statement: Price Intelligence Is a Polling-Dominant Event System
Frames the product as four planes: acquisition, truth, evaluation, and engagement, and separates verified price movement from raw scraped noise.
Problem statement
Design a price watching tool where users add products to a watchlist, set a target threshold, and receive a push or email alert when the current price falls below that threshold. The platform must discover and resolve products across many retailers, re-check prices on a schedule or through direct feeds, store a full history for charting, evaluate watches against every accepted price change, and deliver deduplicated, timely notifications.
This is not a notification app with a small scraping side quest. The dominant workload is acquisition: continuously re-observing tens of millions of product offers across heterogeneous retailer sites under politeness limits, anti-bot pressure, and layout drift. The dominant correctness problem is truth: a scraped price can be wrong because of a layout change, a currency mix-up, a marketplace third-party offer, a lightning-deal artifact, or a retailer glitch. If the platform alerts ten thousand users about a price that was actually a parse error, trust dies immediately and cannot be bought back. Therefore the design separates raw observations from accepted price points, and every downstream consumer, the evaluation engine, the chart API, the analytics warehouse, reads only from the validated canonical ledger.
Why the problem is distinctive
A generic e-commerce backend retries a failed write. A price watcher must also handle the case where the world itself is ambiguous: the retailer now shows a lower price only for logged-in users, only in one region, only with a coupon, or only for one variant. The answer therefore separates price movement from price truth. Price movement is an eventually consistent stream of observations: fetched, extracted, normalized, compared. Price truth is an invariant: a canonical price point enters the ledger only after source validation, confidence scoring, currency and unit normalization, and an anomaly gate. If evidence is contradictory or low-confidence, the correct response is quarantine and stale-label serving, not optimistic fan-out.
The problem requires user-defined thresholds, regular or event-driven price updates, alerting via push or email, historical charts, efficient diff-based updates, caching of current prices, and either scraping or direct feed integration. The section rhythm, visual vocabulary, recap contract, quiz placement, multi-cloud diagrams, and Interview Buddy rules follow the attached orchestrator and core spec.
Public operating baseline versus design assumptions
Public evidence establishes that this category is operationally real and commercially large. PayPal acquired Honey in January 2020 for approximately four billion dollars, with Honey reporting around 17 million monthly active users at the time of the deal; Honey's model pairs a browser extension with a backend price and coupon graph. CamelCamelCamel has tracked Amazon price history since 2008 and built a durable alerts business on top of it. Keepa sells API access to Amazon price and sales-rank history and publishes coverage in the billions of tracked listings across Amazon marketplaces. Profitero, a B2B price intelligence vendor, publicly claims daily tracking on the order of hundreds of millions of products across thousands of retailer sites. These are company-stated public figures, not requirements for our fictional system.
For capacity planning, this answer explicitly assumes a mature consumer price-tracking network with 4 million registered users, 800 thousand monthly active users, 250 thousand daily active users, 12 million active watches over 2.5 million distinct watched products, a 20 million offer catalog across 500 retailer domains, and approximately 51 million price checks per day with a five-times event peak on deal days. Unless a number is tied to a citation, it is a stated design assumption, target, budget, or illustrative threshold, not a claim about any company's private architecture.
The four architectural planes
- Acquisition plane: crawl scheduler, fetch workers, feed and webhook ingestion, politeness enforcement, retry and quarantine.
- Truth plane: extraction, normalization, anomaly gating, canonical price ledger, evidence storage.
- Evaluation plane: watch index, threshold matching, alert candidate generation, deduplication.
- Engagement plane: alert delivery across push, email, and in-app channels; price history charts; watch management UX.
A strong interview answer keeps these planes separate. It allows the engagement plane to degrade without slowing acquisition, and it allows the evaluation plane to catch up after lag without ever reading unvalidated prices from the acquisition plane.
Key Highlights
- •The dominant workload is acquisition: tens of millions of checks per day under per-domain politeness limits, not user request QPS.
- •Separate raw observations from accepted price points; every consumer reads only the validated canonical ledger.
- •A wrong alert is worse than a late alert: anomaly gating and quarantine fail closed for fan-out.
- •The four planes are acquisition, truth, evaluation, and engagement.
- •Public company figures provide context; every uncited scale or SLO in this answer is an explicit design assumption.
Section Rescue Kit
Buzzwords to use:
Safe statements:
- "I will separate price movement, which is an eventual stream, from price truth, which is a validated ledger entry."
- "Before selecting databases, let me define which decisions happen in acquisition, truth, evaluation, and engagement."