Problem Statement: AR Try-On Is a Rendering + Trust Problem, Not a CRUD App
Frames virtual try-on as an edge-first realtime rendering system wrapped around a commerce funnel, where biometric privacy and device fragmentation dominate.
Problem statement
Design a virtual try-on and AR commerce platform that lets shoppers superimpose products onto their own camera view or body in real time — cosmetics on the face, glasses on the head, shoes on the feet, apparel on the body, or furniture in the room — and convert that experience into an add-to-cart action. The platform must serve 3D or shader-based product representations, run face/body tracking and alignment, optionally estimate the shopper's dimensions for fit, measure engagement, and feed the normal commerce funnel.
The first thing to internalize is that this is not a request/response CRUD backend. The defining loop — camera capture, face/body landmark detection, pose estimation, asset rendering, and compositing — must run at 30 to 60 frames per second with a motion-to-photon latency the human visual system tolerates. You cannot put a synchronous cloud RPC in that loop. A 300 ms round trip to an API per frame is 3 fps, which is unusable and would also stream raw biometric video to the cloud, a privacy catastrophe.
Therefore the architecture splits into two planes with very different physics. The rendering plane is on-device: it owns the camera, the trackers, the GPU, and the realtime loop, and it must stay correct and safe with no network. The commerce and asset plane is cloud: it owns the product catalog, the 3D asset pipeline, sizing intelligence, session analytics, cart, and checkout. The cloud delivers immutable, versioned asset bundles ahead of time and consumes low-rate semantic telemetry; it never sits inside the per-frame path.
Why the problem is distinctive
A product-detail backend can retry a lookup. A face-tracking loop cannot retry a dropped frame without visible judder. A food-delivery system can tolerate stale location for a few seconds. An AR overlay that drifts off the user's face by ten pixels looks broken instantly. So the design separates commerce correctness (eventually-consistent analytics, retryable add-to-cart) from perceptual correctness (frame-rate, alignment, photorealism) and from biometric trust (raw face data never leaves the device without explicit, revocable consent).
Three constraints make this category harder than a generic e-commerce feature. First, device fragmentation: an iPhone with a TrueDepth depth camera, a mid-tier Android with a monocular camera, and a laptop browser with only a webcam have wildly different tracking fidelity and render budgets. Second, asset supply: millions of SKUs exist, but almost none have production-ready 3D assets; the platform needs a pipeline to generate, optimize, and version them. Third, biometric regulation: face geometry is special-category data under GDPR and biometric data under laws like Illinois' BIPA, so the default posture must be on-device processing with nothing sensitive uploaded.
Public operating baseline versus design assumptions
Public evidence shows the category is real. Snap reports that over 250 million people engage with AR every day on its camera platform. Google shipped virtual try-on for lipstick in Google Shopping and later foundation try-on powered by an AI model trained across diverse skin tones, using MediaPipe face landmarks. Amazon launched Virtual Try-On for shoes and then eyeglasses in its shopping app. Warby Parker has offered an ARKit-based glasses try-on that uses the TrueDepth camera to measure face width for frame fit since 2017. These are cited company signals, not requirements for our fictional system.
For capacity planning, this answer explicitly assumes a mature fashion/beauty retailer with 10 million MAU, 2 million DAU, 400,000 try-on sessions per day average, and a 5× event peak. Unless a number is tied to a citation, it is a stated design assumption, target, or budget — not a claim about any company's private architecture.
The four architectural planes
- Rendering plane (on-device): camera capture, face/body tracking, pose estimation, lighting estimation, asset loading, GPU shading, compositing, and a local performance governor.
- Asset plane (edge + cloud): 3D asset ingestion, optimization, LOD generation, packaging, signing, CDN distribution, and version manifests.
- Commerce plane (cloud): catalog, sizing and fit engine, recommendations, cart, checkout, and the try-on-to-purchase funnel.
- Trust plane (cross-cutting): consent, biometric minimization, retention, accessibility, and audit.
A strong interview answer keeps these planes separate. It lets the commerce plane degrade without breaking the render loop, and it lets the asset pipeline evolve models and meshes without silently changing what a validated session renders.
Key Highlights
- •The realtime render loop runs entirely on-device; the cloud is never inside the per-frame path.
- •Separate perceptual correctness (frame rate, alignment) from commerce correctness (retryable cart writes).
- •Raw face/body imagery is biometric data; the default is on-device processing with nothing sensitive uploaded.
- •Four planes: rendering, asset, commerce, and trust. Each degrades independently.
- •Every uncited scale number here is an explicit design assumption, labeled as such.
Section Rescue Kit
Buzzwords to use:
Safe statements:
- "I will separate commerce correctness from perceptual correctness before choosing any service."
- "Let me define which decisions run on-device, which run in the cloud, and which require explicit consent."