Design a Warehouse Robot Coordination System

Hard45 min
1 / 30
understanding•10 min read

Problem Statement: A Multi-Robot Coordination Platform

Frames the product as an edge-first fleet orchestrator that couples a real-time digital twin of the warehouse with durable order fulfillment.

Problem statement

Design a warehouse robot coordination system that orchestrates a fleet of autonomous mobile robots (AMRs) to retrieve items from storage, move them to pick or pack stations, avoid collisions, and do so across hundreds of machines simultaneously. The platform must assign tasks, plan collision-free paths over a warehouse grid, track every robot's live pose and battery, integrate with the order queue for pick priority, and recover gracefully when a robot fails.

This is not a simple job scheduler. A robot can lose Wi-Fi behind a tall racking aisle while another robot is converging on the same grid cell at 1.5 m/s. Therefore the immediate safety loop—local obstacle detection, velocity limiting, emergency stop, and cell-reservation enforcement—must run on-device and stay correct without the orchestrator. The orchestrator owns fleet-wide task assignment, path deconfliction, battery scheduling, order prioritization, digital-twin state, and analytics; it must never be the only component able to prevent a collision.

Why the problem is distinctive

A task queue can retry a job. Two robots cannot retry a collision. The design therefore separates fulfillment progress from motion safety. Fulfillment is an eventually progressing workflow: queued, assigned, picking, transporting, delivered, or recovered. Motion safety is an invariant: a robot may move only when its local controller and the grid-reservation protocol together confirm the target cell is clear and reachable. If evidence is stale or contradictory, the safe response is to slow or stop—not to wait for a cloud RPC.

The problem requires robot task assignment, collision avoidance with real-time path adjustment, order-queue integration for pick priority, and telemetry for robot status, battery, and speed. It also calls out low latency, scale to hundreds of robots per warehouse, fault tolerance, and possible edge computing for local real-time decisions.

Public operating baseline versus design assumptions

Public evidence establishes that the category is operationally real. Amazon reports more than 750,000 robots working across its fulfillment network and describes Proteus as its first fully autonomous mobile robot that navigates among employees. Ocado reports hundreds of bots swarming over a single fulfillment grid, coordinated to assemble grocery orders in minutes. Locus Robotics reports tens of thousands of AMRs deployed with more than a billion units picked. These are cited company figures, not requirements for our fictional system.

For capacity planning, this answer explicitly assumes a mature operator with 50 warehouses, 500 active robots per warehouse, 25,000 robots registered globally, and 2,000 pick tasks per warehouse per hour at peak. Unless a number is tied to a citation, it is a stated design assumption, target, budget, or illustrative threshold—not a claim about any company's private architecture.

The four architectural planes

  1. Safety plane: deterministic local controller, emergency stop, cell-reservation enforcement, health checks, and minimal-risk state.
  2. Autonomy plane: localization on the warehouse map, obstacle detection, local path following, and recovery behaviors.
  3. Orchestration plane: task assignment, path deconfliction, order integration, battery and charging scheduling, and fleet state.
  4. Learning and ops plane: telemetry, digital twin, simulation, heatmaps, incident evidence, and policy evolution.

A strong interview answer keeps these planes separate. It allows the orchestration plane to degrade without weakening the safety plane, and it permits the learning plane to improve routing heuristics without silently changing the validated safety envelope.

Key Highlights

  • •The orchestrator may optimize tasks, but each robot preserves motion safety locally during total connectivity loss.
  • •Model pick progress as a durable workflow and cell occupancy as a continuously evaluated reservation invariant.
  • •Public fleet figures provide context; every uncited scale or SLO in this answer is an explicit design assumption.
  • •The architecture has four planes: safety, autonomy, orchestration, and learning/ops.
  • •A safe stop is a successful safety outcome even when it becomes a failed or delayed pick.
Lead With the Safety Boundary
State in the first two minutes that collision avoidance and emergency stopping are local and independent of orchestrator availability. This instantly distinguishes a robotics architecture from a generic task queue.
Do Not Draw a Cloud-Controlled Toy
A design where every obstacle decision or brake command depends on a round trip to the orchestrator is unsafe under ordinary Wi-Fi loss behind racking and will fail a serious interview.

Section Rescue Kit

Buzzwords to use:

Cell Reservation ProtocolMinimal-Risk Condition

Safe statements:

  • "I will separate pick completion from motion safety: the former may retry, while the latter must fail closed locally."
  • "Before selecting services, let me define which decisions are allowed on-device, in the orchestrator, and only with human approval."
Design a Warehouse Robot Coordination System - System Design | WinJob | WinJob