Design an RPA System with Vision + LLM for Agents

Hard60 min
1 / 25
understanding•8 min read

Problem Statement: The Fragility of Traditional RPA

Frames the transition from brittle DOM-based automation to resilient, vision-powered agentic workflows.

Problem statement

Design an enterprise Robotic Process Automation (RPA) platform that executes complex business workflows across heterogeneous web and desktop applications. Unlike traditional RPA, which relies on brittle Document Object Model (DOM) selectors and pixel-coordinate mapping, this system must integrate Computer Vision (CV) and Multimodal Large Language Models (LLMs) to interpret ambiguous user interfaces, handle dynamic pop-ups, and recover from layout changes autonomously.

Traditional RPA breaks when a web application updates its CSS classes, changes a button's text, or introduces an unexpected modal dialog. The maintenance burden of 'healing' these broken selectors consumes up to 30% of an RPA center of excellence's budget. By treating the screen as a visual canvas rather than a structured tree, and using an LLM as a reasoning engine, the agent can understand the intent of a UI element (e.g., 'the blue button that says Submit') rather than its underlying HTML ID.

Why the problem is distinctive

A standard backend API integration retries a 500 error. An RPA agent interacting with a legacy desktop application cannot retry a misclick that accidentally deletes a production database record. The design therefore separates workflow orchestration from visual grounding and action execution. Workflow orchestration is a durable, eventually consistent state machine managed by the cloud. Visual grounding and action execution are latency-sensitive, safety-critical loops that must execute locally on the agent node or in a tightly controlled sandboxed environment.

The platform must schedule processes, manage credential vaults, stream screen captures to a vision pipeline, invoke LLMs for semantic reasoning, execute mouse/keyboard actions, and escalate to a Human-in-the-Loop (HITL) when confidence falls below a safety threshold.

Public operating baseline versus design assumptions

Public evidence establishes the scale of the RPA market. UiPath reports processing billions of transactions annually across thousands of enterprise customers, with their Computer Vision service specifically designed to bypass DOM fragility. Adept AI and Anthropic have recently demonstrated 'Computer Use' paradigms where multimodal models directly output mouse coordinates and keystrokes based on visual prompts. These represent the cutting edge of the agentic RPA space.

For capacity planning, this answer explicitly assumes a mature enterprise deployment with 100,000 concurrent software robots (agents), 50 million UI interactions per day, and 5 million multimodal LLM reasoning calls per day. Unless a number is tied to a citation, it is a stated design assumption.

The three architectural planes

  1. Control Plane: Process definition, scheduling, credential management, audit logging, and fleet orchestration.
  2. Vision & Reasoning Plane: Screen capture, OCR, object detection (Set-of-Mark), semantic caching, and LLM tool-use (ReAct) loops.
  3. Execution Plane: OS-level input injection, sandboxing, distributed session locking, and local safety guardrails.

Key Highlights

  • •Traditional DOM-based RPA is brittle; Vision+LLM RPA treats the screen as a semantic canvas.
  • •Separate durable workflow orchestration from latency-sensitive visual grounding and action execution.
  • •Public figures from UiPath and Anthropic provide context; every uncited scale number is an explicit design assumption.
  • •The architecture has three planes: Control, Vision/Reasoning, and Execution.
Lead With the Fragility Problem
State in the first two minutes that DOM selectors break when CSS changes. This instantly justifies the need for a Computer Vision and LLM pipeline, distinguishing your design from a basic scripting engine.
Do Not Route Every Click Through the Cloud
Sending every screenshot to a cloud LLM for every mouse movement will result in unacceptable latency and massive token costs. Vision grounding must be hierarchical and heavily cached.

Section Rescue Kit

Buzzwords to use:

Visual GroundingSelf-Healing Automation

Safe statements:

  • "I will separate the durable workflow state from the real-time visual grounding loop."
  • "Before selecting models, let me define the latency and cost budgets for screen interpretation."
Design an RPA System with Vision + LLM for Agents - System Design | WinJob | WinJob