Problem Statement: A Continual Perception Platform, Not a Batch Classifier
Frames shelf recognition as an always-on cyber-physical ML system with four planes: perception, knowledge, operations, and learning.
Problem statement
Design a continual visual recognition platform for retail shelves. Fixed cameras (or an equivalent capture fleet) observe store shelves and the system must detect out-of-stock facings, low stock, misplaced products, missing or wrong price labels, and planogram violations. It must alert store staff with enough context to act, close the loop when staff fix the shelf, and keep improving as the product catalog, packaging, promotions, and store layouts change week after week.
The word continual is the design driver. A retail assortment is not a static classification dataset. Chains introduce hundreds of new SKUs every week, suppliers redesign packaging without notice, seasonal items appear and disappear, planograms reset overnight, and lighting changes with the season. A model that was accurate at launch silently degrades unless the architecture includes an explicit data flywheel: observation, uncertainty capture, prioritized labeling, versioned retraining, per-cohort evaluation, and guarded rollout.
This is also not a cloud-only inference problem. A mid-size supermarket with 60 shelf cameras at one frame per second generates 60 frames per second locally. Uploading compressed frames from 2,000 stores would require gigabytes per second of WAN egress, which is economically and physically impossible on business broadband. Detection therefore runs at a store-level edge node, and only semantic shelf-state events, alerts, and small evidence clips travel to the cloud.
Why the problem is distinctive
A generic image-classification backend can tolerate stale predictions. A shelf system cannot tolerate three correlated failure classes: false out-of-stock alerts that burn staff trust, missed out-of-stocks that lose revenue while dashboards show green, and model regression that hits hundreds of stores at once after a bad release. The design therefore separates three concerns that naive answers merge.
First, perception freshness: what does the shelf look like right now, with quantified uncertainty and occlusion handling. Second, alert correctness: which perceptions become durable, deduplicated, prioritized staff tasks. Third, model lifecycle: how the system learns without forgetting, and how releases are gated by operational evidence rather than benchmark accuracy alone.
The four architectural planes
- Perception plane (store edge): frame acquisition, camera calibration, shelf-region projection, privacy masking, object detection, SKU identification, temporal consensus, and health monitoring.
- Knowledge plane: product catalog and embeddings, planogram versions, shelf-region maps, camera-to-shelf calibration, and ODD-style operating constraints per store format.
- Operations plane: shelf-state projection, alert engine, staff tasking, proof of resolution, store and HQ dashboards, OSA and compliance reporting, and CPG data products.
- Learning plane: drift monitoring, hard-example mining, labeling queues, dataset and model versioning, evaluation gates, canary rollout, rollback, and incident evidence.
A strong interview answer keeps these planes separate. The operations plane may degrade during a WAN outage while the perception plane keeps detecting locally, and the learning plane may improve models without silently changing what is live in stores.
Public operating baseline versus design assumptions
Public evidence shows the category is real. Trax markets fixed-camera shelf monitoring (Trax Watch) plus image recognition used by CPG field teams. Simbe Robotics operates the Tally autonomous shelf-scanning robot with public retail deployments such as Schnucks Markets. Pensa Systems describes drone and camera-based shelf scanning. Focal Systems sells fixed-camera OSA alerting. Amazon publicly built multi-camera store perception for Just Walk Out and Dash Cart. These are cited as directional context, not as numbers for our fictional system.
For capacity planning, this answer explicitly assumes a mature chain deployment: 2,000 stores, 60 cameras per store on average, 120,000 cameras registered, about 100,000 active during opening hours, one frame per second per camera, and a 5x event peak around promotions, holidays, and store resets. Unless tied to a named public source, every number here is a stated design assumption.
Key Highlights
- •Continual distribution shift, not one-time training, is the core design driver.
- •Edge inference is mandatory: 60 cameras at 1 fps per store makes raw frame upload infeasible.
- •Separate perception freshness, alert correctness, and model lifecycle into distinct planes.
- •Four planes: perception, knowledge, operations, and learning.
- •A false out-of-stock alert is a trust cost; a missed out-of-stock is a revenue cost; both are measurable.
Section Rescue Kit
Buzzwords to use:
Safe statements:
- "Let me separate what the model sees from what the business acts on before drawing services."
- "I will treat model change as a constant, not an exception, and design the release path first."