Design a Video Analytics System for Traffic Monitoring with ML

Medium45 min
1 / 30
understanding•10 min read

Problem Statement: A City-Wide Vision Pipeline, Not a Single Camera

Frames the system as a multi-camera edge-to-cloud vision platform where detection, metric computation, and alerting must survive partial camera and network failure.

Problem statement

Design a video analytics platform for city traffic monitoring that ingests live feeds from hundreds to thousands of intersection and highway cameras, runs ML object detection and tracking on each stream, derives traffic metrics (vehicle count, average speed, queue length, occupancy, congestion index), detects anomalies such as collisions, wrong-way drivers, stopped vehicles, and major slowdowns, and raises near-real-time alerts that can auto-dispatch to traffic operators or emergency services. It also stores derived metrics and selected clips for historical trend analysis and planning.

This is not a single-camera computer-vision exercise. The defining challenge is concurrency across a large camera fleet. A pilot with five cameras is a demo; a production deployment with 2,000 cameras is a distributed-systems problem where per-stream inference cost, bandwidth, partial outages, and model rollout across heterogeneous hardware dominate the design. The brief is explicit that competitors stop at "traffic cam" and ignore multi-camera city-scale coverage; that gap is exactly where this answer must go deep.

Why the problem is distinctive

A generic video-processing backend can drop a frame and retry. A traffic-safety pipeline cannot silently lose an accident. The design therefore separates detection freshness from alert correctness. Metrics like vehicle counts and speed are eventually consistent, lossy, and freshness-bounded — a 5-second-old count is still useful. But an accident alert is a consequential event: it must be deduplicated, confirmed, and delivered even when the camera's region is degraded. We model metrics as a high-rate lossy stream and alerts as a durable, exactly-once-ish workflow.

The other distinctive axis is where inference runs. Pushing 2,000 raw 1080p streams to a central data center is a bandwidth and cost disaster. The credible architecture pushes detection to the edge (on-camera or edge nodes) so that only compact structured events (bounding boxes, track IDs, counts, speeds) cross the network, with raw or clipped video uploaded selectively for evidence. This is the same edge-first lesson as autonomous fleets and IoT: the cloud optimizes and aggregates, the edge perceives.

Public operating baseline versus design assumptions

Public evidence establishes that the category is operationally real. NVIDIA's Metropolis/DeepStream platform documents multi-stream video analytics on GPUs, with DeepStream able to decode and infer on dozens of concurrent 1080p streams per accelerator. Miovision builds dedicated traffic-counting cameras and publishes intersection analytics. Iteris has shipped video vehicle detection for signal control for decades. Alibaba's CityBrain in Hangzhou is a widely cited city-scale traffic AI deployment. These are cited as category context, not as our design targets.

For capacity planning this answer explicitly assumes a mid-size city deployment: 2,000 cameras registered, 1,600 simultaneously streaming, 1080p at 15 fps for analytics, H.265 at an average of 3 Mbps per stream, and a 5x incident peak. Unless a number is tied to a citation, it is a stated design assumption, target, or budget — never a claim about any vendor's private architecture.

The four architectural planes

  1. Ingestion plane: camera connectivity, RTSP/RTMP/GB28181 pull or push, edge decode, health, and adaptive frame sampling.
  2. Inference plane: object detection, multi-object tracking, per-camera metric extraction, running on edge GPUs or regional inference clusters.
  3. Analytics plane: stream aggregation into intersection, corridor, and city-level metrics; time-series storage; anomaly and incident detection.
  4. Action & learning plane: alerting and dispatch, dashboards, historical trend analysis, model evaluation, dataset curation, and governed model rollout.

A strong answer keeps these planes separate. It lets the analytics plane lag without breaking detection, and it lets the learning plane improve models without silently changing a validated detection envelope.

Key Highlights

  • •The hard problem is concurrency across thousands of cameras, not single-stream accuracy.
  • •Detection runs at the edge; only structured events cross the network by default.
  • •Metrics are lossy and freshness-bounded; accident alerts are durable, deduplicated workflows.
  • •The architecture has four planes: ingestion, inference, analytics, and action/learning.
  • •Every uncited scale or SLO in this answer is an explicit design assumption.
Lead With Fleet Concurrency
State in the first two minutes that the design is driven by thousands of concurrent cameras, not by one model's accuracy. This instantly separates a city-scale architecture from a computer-vision homework answer.
Do Not Centralize Raw Video
A design that ships 2,000 raw 1080p streams to one data center ignores bandwidth and cost reality and will fail a serious interview. Ship structured events; upload clips selectively.

Section Rescue Kit

Buzzwords to use:

Edge InferenceStructured Event

Safe statements:

  • "I will separate detection freshness from alert correctness: counts may be lossy, accidents may not."
  • "Before choosing services, let me define what runs at the edge versus in the cloud."
Design a Video Analytics System for Traffic Monitoring with ML - System Design | WinJob | WinJob