Design a Self-Driving Car Perception & Planning Stack

Hard60 min
1 / 30
understanding•11 min read

Problem Statement: A Deadline-Bound Cyber-Physical Perception and Planning Stack

Frames the stack as a real-time embedded AI system whose correctness is measured in milliseconds and meters, not requests per second.

Problem statement

Design the perception and planning stack of a Level-4 autonomous vehicle: ingest heterogeneous sensors (cameras, lidar, radar, IMU, GNSS, wheel odometry), fuse them into a coherent, uncertainty-aware world model, predict the future motion of every relevant agent, and produce a dynamically feasible, collision-free trajectory that a controller executes at 100 Hz. The stack must survive partial sensor failure, compute overrun, and sensor disagreement, and it must degrade to a defined minimal-risk behavior instead of guessing.

This is not a cloud ML serving problem. A perception cycle that misses its deadline by 80 ms at 50 km/h means 1.1 meters of blind travel. Therefore the authoritative loop—synchronization, detection, tracking, prediction, planning, control—runs on-vehicle on deterministic compute with worst-case execution time budgets. The cloud owns the slow loop: data engine, labeling, training, simulation, model release, maps, fleet telemetry, and remote assistance. The cloud may improve the driver tomorrow; it may never be required to brake today.

Why the problem is distinctive

Three properties separate this design from a generic ML platform. First, concurrency and time: eight camera streams, three lidars, and five radars must be temporally aligned to single-digit milliseconds before fusion, because a 100 ms skew between a camera frame and a lidar sweep produces a phantom obstacle or a missed pedestrian. Second, uncertainty as a first-class citizen: every detection carries covariance, every prediction carries a probability over multiple futures, and the planner must consume distributions, not point estimates. Third, certified failure behavior: functional-safety standards require that each hazardous failure mode has a detected, bounded, tested response—degraded perception modes, minimal-risk maneuvers, and independent monitoring—rather than a retry.

Public operating baseline versus design assumptions

Public evidence shows the category is real. Waymo has reported tens of millions of autonomous public-road miles and billions of simulated miles in its safety materials; Tesla has reported billions of cumulative FSD (Supervised) miles in investor disclosures; Baidu Apollo operates Apollo Go robotaxi services and publishes its open-source stack; NVIDIA publishes DRIVE Orin at 254 TOPS and DRIVE Thor at 2,000 TOPS as reference compute. These are cited company figures for context. Unless a number is tied to such a citation, every fleet size, data rate, latency budget, and storage figure in this answer is an explicitly stated design assumption for back-of-envelope sizing.

The five architectural planes

  1. Sensing plane: sensors, time synchronization, calibration health, raw ingest, ISP and preprocessing.
  2. World-model plane: detection, free space, occupancy, tracking, localization context, uncertainty propagation.
  3. Decision plane: prediction, behavior selection, motion planning, trajectory validation.
  4. Actuation and safety plane: control, actuator interface, independent safety monitor, minimal-risk maneuver executor.
  5. Fleet-learning plane: selective logging, data engine, labeling, training, simulation, shadow mode, governed model release.

A strong interview answer keeps these planes separate, states which decisions are allowed to depend on the cloud (none in the immediate loop), and treats every deadline as a design artifact with an owner and a measured budget.

Key Highlights

  • •The authoritative perception-to-control loop is on-vehicle and deadline-bound; the cloud owns only the slow learning loop.
  • •Temporal alignment across sensors is a correctness requirement: 100 ms skew at 50 km/h is 1.4 m of positional error.
  • •Uncertainty (covariance, multi-modal futures) flows from perception through prediction into planning constraints.
  • •Five planes: sensing, world model, decision, actuation/safety, fleet learning.
  • •Company figures (Waymo miles, Tesla FSD miles, NVIDIA TOPS) are context; all other numbers are labeled assumptions.
Lead With Deadlines, Not Models
Open by converting speed into distance-per-millisecond: at 50 km/h the vehicle travels 13.9 m per second, so a 100 ms perception delay is 1.4 m of unobserved motion. This instantly signals that you design real-time systems, not dashboards.
Do Not Put Inference in the Cloud
A design that streams raw sensor data to a cloud model for detection fails on latency, bandwidth, and availability simultaneously. Cellular RTT alone (30-80 ms) plus upload of hundreds of MB/s is infeasible; braking must never depend on it.

Section Rescue Kit

Buzzwords to use:

Operational Design DomainGlass-to-Actuation Latency

Safe statements:

  • "Let me separate the on-vehicle real-time loop from the cloud learning loop before choosing any technology."
  • "Before discussing models, I will state the deadline each stage owns and what happens when it misses."
Design a Self-Driving Car Perception & Planning Stack - System Design | WinJob | WinJob