Design a Workflow of LLM Agents for Data Analysis

Medium45 min
1 / 31
understanding•10 min read

Problem Statement: A Durable Multi-Agent Pipeline, Not a Chatbot

Frames the system as a long-running, artifact-passing workflow of specialized LLM agents with partial-failure semantics and human oversight.

Problem statement

Design a platform that automates end-to-end data analysis using a chain of LLM-based agents. A user uploads a dataset and states a goal in natural language. The system then runs a pipeline of specialized agents: a Cleaning agent that profiles the data, fixes schemas, handles missing values, and emits a cleaned dataset; an EDA agent that forms hypotheses, writes and executes analysis code in a sandbox, and produces charts and summary statistics; a Modeling agent that engineers features, fits and evaluates candidate models, and emits model artifacts with metrics; and an Interpretation agent that writes a narrative report grounded in the intermediate artifacts, with caveats about leakage, imbalance, and statistical validity. Agents must pass intermediate artifacts — cleaned CSV or Parquet subsets, summary statistics, chart objects, model binaries — from one stage to the next, must tolerate partial errors such as a misclassified column or a model that fails to converge, and must expose every step for user inspection and final validation.

This is not a chatbot with a Python tool. A chat session is short-lived, single-threaded, and forgiving: if a step fails, the user rephrases. An agentic data-analysis pipeline is a long-running distributed workflow that may execute 15 to 40 LLM calls and 20 to 60 sandboxed code executions over 5 to 30 minutes, spending real money per run, touching data that may contain PII, and producing artifacts that downstream decisions depend on. The design therefore splits cleanly into four planes, and the interview answer that wins is the one that keeps them separate.

The four architectural planes

  1. Control plane: a durable workflow orchestrator that owns the pipeline state machine, checkpoints every stage, schedules agent tasks, enforces budgets, and resumes after any crash without re-executing completed work.
  2. Execution plane: ephemeral, network-isolated code sandboxes where agent-generated Python actually runs, with CPU, memory, disk, and wall-clock caps, and with artifacts written to a content-addressed store rather than a local disk.
  3. Intelligence plane: the LLM gateway — provider routing, retries, token accounting, prompt and tool schemas, caching, and per-workflow budget enforcement. This is where Anthropic's published observation that multi-agent systems can consume roughly 15x the tokens of a simple chat becomes an architectural problem, not a trivia fact.
  4. Oversight plane: human-in-the-loop gates, step inspection UIs, approval leases, audit trails, and the final validation step where a human signs off on the report before it is published.

Why this problem is distinctive

Three properties separate this from ordinary backend design. First, non-determinism is in the critical path: the same stage can produce syntactically valid but semantically wrong code, so every stage needs programmatic validation gates — schema checks on cleaned data, metric thresholds on models — not just exit-code checks. Second, artifacts are the state: the workflow's true state is not a row in a database, it is a graph of immutable artifacts (dataset v2, stats.json, model.pkl, report.md) with lineage edges. Losing or mis-referencing one artifact corrupts every downstream stage. Third, the failure unit is a stage, not the workflow: a misclassified column in cleaning should be detectable at EDA, fixable by re-running cleaning with corrective feedback, and recoverable without restarting the entire run. That is exactly the brief's requirement to 'handle partial errors or misclassification.'

Public baseline versus design assumptions

The category is real and operating at scale. LangChain's State of AI Agents survey, published in December 2024 from over 1,300 respondents across industries, found that 51% of organizations already had agents in production, with reliability and cost the top blockers. Anthropic's engineering blog on its multi-agent research system describes an orchestrator-worker pattern where a lead agent spawns parallel subagents, and reports that multi-agent configurations consume roughly 15x the tokens of chat while outperforming single-agent setups on parallelizable research work. OpenAI has run sandboxed Python code execution for hundreds of millions of ChatGPT users. These are cited context figures. Every scale number used for capacity planning in this answer — workflow counts, token budgets, sandbox concurrency — is an explicitly stated design assumption, labeled where it appears.

Key Highlights

  • •Four planes: durable control plane, sandboxed execution plane, LLM intelligence plane, human oversight plane.
  • •Agents pass typed, versioned artifacts — cleaned datasets, stats, models, reports — not chat transcripts.
  • •Stage-level validation gates catch semantically wrong but syntactically valid agent output.
  • •The failure unit is one stage plus its artifacts, so partial errors are repairable without full restart.
  • •Anthropic publicly notes multi-agent systems can consume ~15x chat token volume, making budgeting a first-class design concern.
Lead With Durability
State in the first two minutes that this is a long-running durable workflow with artifact lineage, not a chat feature. That single framing separates staff-level answers from junior ones.
Do Not Draw One Big Agent Loop
A single free-form agent loop with no stage boundaries cannot checkpoint, cannot budget per stage, cannot be inspected step by step, and cannot recover from a mid-pipeline failure. The brief explicitly requires distinct agents with distinct skills.

Section Rescue Kit

Buzzwords to use:

Orchestrator-Worker PatternArtifact Lineage

Safe statements:

  • "I will separate workflow durability from LLM nondeterminism before choosing any storage or compute."
  • "The design question is not which model to call, it is how stages pass validated artifacts and recover from partial failure."
Design a Workflow of LLM Agents for Data Analysis - System Design | WinJob | WinJob