Problem Statement: A Money-Moving Decision in Under 100 Milliseconds
Frames RTB as a latency-bound, money-critical ML serving problem rather than a generic ads CRUD system.
Problem statement
Design the real-time bidding (RTB) core of a demand-side platform (DSP). When a user opens a page or app, the publisher's supply-side platform (SSP) or ad exchange broadcasts an OpenRTB bid request describing the user, device, site, app, ad slot, floor price, and privacy signals. Our bidder must decide, within the exchange's deadline, whether to bid, which campaign and creative to bid for, how much to pay, and return a well-formed bid response. The IAB Tech Lab's OpenRTB specification and Google's published Ad Exchange documentation both converge on a bid window of roughly 100 milliseconds, and anything slower is silently dropped: a late bid is not a worse bid, it is a non-existent bid.
The decision is an ML pipeline under a hard real-time constraint. A CTR or conversion-rate model scores the opportunity using user features, context features, and campaign features; a bid price is derived from the predicted value, the auction type, and the campaign's budget state; a pacing controller decides whether the campaign is allowed to spend right now; and a privacy gate decides whether we are legally permitted to bid at all. Every one of these steps competes for the same 100 milliseconds, and every one of them can fail independently.
Why this problem is distinctive
Three properties separate RTB from ordinary ML serving. First, the workload is enormous and mostly worthless: a large DSP receives hundreds of thousands to millions of bid requests per second, and a correct design declines the overwhelming majority of them cheaply before spending feature-fetch or inference cost. Second, the output moves real money immediately: a bug in budget enforcement does not produce a bad recommendation, it produces overspend measured in dollars per minute, so the money path needs stronger consistency than the prediction path. Third, the labels arrive late and biased: clicks trickle in over minutes, conversions over days, and both are conditioned on having won an auction at a particular price, so the training loop must correct for auction and throttling bias or the model silently learns our own bidding policy instead of user behavior.
The four architectural planes
- Exchange edge plane: per-point-of-presence bid endpoints, request parsing, pre-bid filtering, deadline enforcement, and response serialization close to the exchange.
- Decision plane: feature retrieval, CTR/CVR inference, bid pricing, bid shading, pacing, and budget checks.
- Money plane: campaign budgets, spend ledgers, win notices, billing reconciliation, loss reasons, and fraud-filtered chargeable events.
- Learning plane: sampled request logs, joined win/click/conversion feedback, feature pipelines, model training, evaluation, shadowing, and guarded promotion.
A strong interview answer keeps these planes separate: the decision plane may degrade to a simpler model, but the money plane must never degrade into optimistic spending, and the learning plane must never push a model into the decision plane without evidence gates.
Public scale context versus design assumptions
Public filings establish the economic scale: Alphabet reported advertising revenue of about 237.9 billion USD for fiscal 2023 and Meta about 131.9 billion USD, and IAB/PwC measurement puts programmatic as the dominant transaction method for US display advertising. Company-described DSP engineering talks, for example from The Trade Desk and Criteo, consistently describe bid-request intake in the millions per second during peak. Those are context, not requirements. For capacity planning this answer explicitly assumes a large DSP with 1,000,000 bid requests per second at peak, 400,000 per second average, a 10 percent pre-bid filter pass rate, and a 100 ms exchange deadline. Unless a number is tied to a named public source, it is a stated design assumption, target, or budget.
Key Highlights
- •OpenRTB and Google AdX documentation converge on a ~100 ms bid window; a late bid is a lost bid, not a slow bid.
- •Most bid requests must be declined cheaply: scoring 100 percent of traffic is the classic cost and latency mistake.
- •The money path (budgets, wins, billing) needs stronger consistency than the prediction path (features, scores).
- •Labels are delayed and auction-biased, so the training loop needs importance correction, not just more data.
- •Four planes: exchange edge, decision, money, learning; each degrades differently under failure.
Section Rescue Kit
Buzzwords to use:
Safe statements:
- "Let me separate the auction-facing edge from the money and learning planes before choosing any technology."
- "I will label every scale number as either a published figure or an explicit assumption before using it in capacity math."