Design a ChatGPT Plugin for Live Code Testing & Execution

Medium45 min
1 / 30
understanding•10 min read

Problem Statement: An LLM With Hands to Run Code

Frames the plugin as a sandboxed execution fabric wrapped around a chat model, not as a chatbot feature.

Problem statement

Design a ChatGPT plugin that lets users write code snippets inside a conversation, executes them in a secure sandbox, returns the runtime output or error back into the LLM context, and lets the model refine the code across bounded iterations. The problem requires secure container or VM execution, LLM API integration that passes user code and receives run results, session-based snippet versioning, concurrency for many simultaneous users, a sandbox that survives hostile code, and multi-language support.

The distinctive fact about this system is that the LLM is both a producer and a consumer of execution results. The user asks a question, the model emits code, the platform runs it, and stdout, stderr, exit codes, and generated files flow back into the next model call. That creates a closed loop: model -> tool call -> sandbox -> shaped output -> model. Everything in the architecture exists to make that loop safe, bounded, and fast.

Why this is not a chat feature

A plain chat backend retries a failed message. An execution backend cannot retry a side effect blindly, because untrusted code may have written files, exhausted memory, or hung. The design therefore separates three concerns that weaker answers collapse into one:

  1. Conversation state: the ChatGPT thread, owned by OpenAI's platform, visible to us only through the plugin request context.
  2. Execution state: our source of truth. Snippets, versions, runs, outputs, artifacts, sandbox leases.
  3. Model feedback: the trimmed, sanitized projection of execution state that is safe and useful to place back into a prompt.

The LLM must never see raw multi-megabyte stack traces, secrets, or other users' data. The platform must never execute code outside a jailed environment. The user must never wait on a synchronous HTTP call while a 60-second run completes. These three constraints drive every later decision.

Public operating baseline

The category is real and large. OpenAI reported roughly 800 million weekly active ChatGPT users at DevDay in October 2025, and its built-in Code Interpreter (now Advanced Data Analysis) has shipped containerized Python execution to hundreds of millions of users since mid-2023. Practitioner reverse-engineering of Code Interpreter, published by Simon Willison and others, describes a Linux container with roughly 1 GB of memory, an ephemeral working directory, no outbound network access, and a hard execution timeout. E2B, an open-source sandbox provider for AI agents, publicly documents Firecracker microVM sandboxes with sub-150 millisecond startup using snapshot restore. Replit states it serves tens of millions of users running untrusted code daily and has publicly discussed container and VM isolation for user workspaces.

For capacity planning this answer assumes a mature plugin with 50 million weekly-eligible users and 1 million code executions per day. Every uncited number below is an explicit design assumption, target, or budget, not a company measurement.

The four architectural planes

  1. Plugin edge plane: manifest, OpenAPI contract, request validation, tool-call semantics.
  2. Orchestration plane: sessions, snippet versions, execution queue, sandbox pool allocator.
  3. Execution plane: Firecracker microVMs or gVisor containers, per-language runner daemons, resource limits, output capture.
  4. Feedback plane: output shaping, error extraction, artifact signing, iteration budgeting, version history.

A strong answer keeps these planes separate so that a sandbox pool outage degrades the feedback plane gracefully while the conversation plane keeps working.

Key Highlights

  • •The LLM is both producer and consumer of execution results; the architecture is a bounded closed loop.
  • •Separate conversation state (OpenAI-owned), execution state (our source of truth), and model feedback (a sanitized projection).
  • •OpenAI reported about 800M weekly ChatGPT users at DevDay in October 2025; Code Interpreter runs containerized Python with no network access.
  • •Assume 50M weekly-eligible users and 1M executions per day; every uncited number is a stated assumption.
  • •Four planes: plugin edge, orchestration, execution, feedback.
Lead With the Closed Loop
State in the first two minutes that the model both emits code and consumes execution results. Naming the loop (model -> tool call -> sandbox -> shaped output -> model) instantly separates this from a generic REST service design.
Do Not Draw a Synchronous Exec
A design where the ChatGPT HTTP call blocks for 60 seconds while Python runs will fail the interview. Executions must be accepted asynchronously with bounded polling or callbacks.

Section Rescue Kit

Buzzwords to use:

Tool-Use LoopExecution Fabric

Safe statements:

  • "I will separate conversation state, execution state, and model feedback before choosing any storage."
  • "The sandbox must survive hostile code, so I will treat the execution plane as a security boundary first and a compute service second."
Design a ChatGPT Plugin for Live Code Testing & Execution - System Design | WinJob | WinJob