Problem Statement: An Agentic Code Generation and Auto-Verification Platform
Frames the product as a closed-loop generate-test-refine pipeline, not a single LLM call.
Problem statement
Design a GPT-based code generation and testing service that accepts a natural-language prompt describing desired functionality, uses an LLM to produce a code stub or a complete module, automatically compiles and executes a test harness against it, and iteratively refines the code when tests fail until they pass or a budget is exhausted. The platform must support multiple programming languages and frameworks, stream partial results for a responsive developer experience, and never run untrusted generated code outside a hardened sandbox.
This is not a chatbot that returns a string. It is a closed-loop cyber-physical-of-software system: the LLM is a stochastic producer, the test runner is a deterministic verifier, and an orchestrator closes the loop. The distinguishing property is that correctness is measured, not asserted. A response that compiles and passes its own generated tests is categorically different from a plausible-looking snippet, and the entire architecture exists to make that measurement cheap, safe, and fast.
Why the problem is distinctive
A normal web backend can retry a failed request. Here, the "retry" is expensive and open-ended: each refinement iteration costs another full LLM inference (hundreds of milliseconds to many seconds and real dollars) plus another sandboxed build-and-test cycle. A naive design loops forever on an impossible prompt, burning tokens and sandbox-minutes. Therefore the design must separate generation quality (an eventually-improving loop) from resource safety (a hard invariant: bounded iterations, bounded wall-clock, bounded spend, and no sandbox escape).
The problem requires LLM-based generation, a test harness for quick compilation and execution, iterative refinement on failure, multi-language support, low latency, scalability under high developer concurrency, reliability against infinite failing loops, and safety checks for potentially malicious code.
Public operating baseline versus design assumptions
Public evidence establishes that the category is real and large. GitHub reported Copilot surpassed 1 million paying subscribers in 2023 and grew past 1.8 million by mid-2024, and has publicly claimed that in enabled files Copilot can write roughly 46% of the code. Replit has reported tens of millions of users and runs untrusted user code in isolated environments built on Nix and Firecracker-style microVMs. OpenAI's Codex and later GPT-4-class models power many of these assistants, and DeepMind's AlphaCode demonstrated a generate-then-filter strategy for competitive programming. These are cited company figures and product descriptions, not requirements for our fictional system.
For capacity planning this answer explicitly assumes a mature platform with 500,000 registered developers, 60,000 simultaneously active at peak, 2,000,000 generation requests per day, and a five-times event peak. Each request triggers an average of three refinement iterations. Unless a number is tied to a citation, it is a stated design assumption, target, budget, or illustrative threshold—not a claim about any company's private architecture.
The four architectural planes
- Generation plane: prompt construction, context retrieval, LLM inference, streaming, and artifact extraction.
- Verification plane: sandboxed build, test execution, lint/static analysis, and structured result capture.
- Orchestration plane: the generate-test-refine loop, budgets, state machine, and terminal outcomes.
- Fleet-learning plane: telemetry, evaluation, prompt/model A/B testing, incident evidence, and governed release.
A strong interview answer keeps these planes separate. It lets the orchestration plane degrade without weakening the verification plane's isolation guarantees, and it lets the learning plane improve prompts and models without silently changing the validated safety envelope around sandbox execution.
Key Highlights
- •The LLM is a stochastic producer and the test harness is a deterministic verifier; the orchestrator closes the loop.
- •Correctness is measured by execution, not asserted by the model — that is the product's defining property.
- •Separate generation quality (an improving loop) from resource safety (hard budgets and sandbox isolation).
- •Public fleet figures from GitHub and Replit provide context; every uncited scale or SLO here is an explicit assumption.
- •The architecture has four planes: generation, verification, orchestration, and fleet learning.
Section Rescue Kit
Buzzwords to use:
Safe statements:
- "I will separate generation quality, which improves in a loop, from resource safety, which is a hard invariant."
- "Before selecting services, let me define which decisions are made by the model, by the verifier, and only by the orchestrator's budget."