Design a Generative Image/Video Service (Stable Diffusion style)

Hard45 min
1 / 30
understanding•10 min read

Problem Statement: GPU-Bound Generative Inference at Consumer Scale

Frames the product as a GPU-orchestration and async delivery problem, not a simple request/response API.

Problem statement

Design a generative image and video service that accepts text prompts (and optionally reference images), runs a diffusion-based model to produce new images or short video clips, and delivers results to users. The service must orchestrate expensive GPU inference, handle user concurrency spikes, store generated media, and enforce content safety checks before delivery.

This is fundamentally different from a CRUD web service. A single image generation on SDXL takes 4-12 seconds of continuous A100 GPU time. A short video clip via Stable Video Diffusion takes 30-90 seconds. These are not millisecond operations. The system is therefore a long-running job orchestration platform where the compute resource (GPU) is scarce, expensive, and stateful (model weights must be loaded into VRAM before inference begins).

Why this problem is distinctive

A typical web API can add horizontal replicas behind a load balancer. A GPU inference service cannot simply add replicas because:

  1. Model loading latency: Loading SDXL weights (~6.9 GB) into VRAM takes 10-30 seconds from network storage, or 2-5 seconds from local NVMe. A cold worker cannot serve immediately.
  2. VRAM is finite: An A100-40GB can hold roughly 4-6 concurrent SDXL inference slots depending on batch size and precision. Overcommitting causes OOM kills.
  3. Cost asymmetry: An idle A100 costs $2-4/hour whether or not it processes work. Overprovisioning burns cash; underprovisioning creates queue blowout.
  4. Heterogeneous workloads: Image generation (SDXL, 4-12s), image-to-image (similar), and video generation (SVD, 30-90s) have wildly different resource profiles and cannot share a single homogeneous worker pool efficiently.

The problem requires prompt ingestion, GPU scheduling, content filtering, and result delivery. The design must handle partial user prompts, large GPU usage, and emergent bridging between diffusion models and real-time service expectations.

The four architectural planes

  1. Ingestion plane: API gateway, prompt validation, safety pre-screening, job creation, user-facing status.
  2. Orchestration plane: job queue, priority lanes, GPU worker pool management, autoscaling, model registry.
  3. Inference plane: GPU workers, model loading, batch assembly, diffusion execution, result encoding.
  4. Delivery plane: object storage, CDN, result notification, moderation post-check, user gallery.

A strong answer keeps these planes separate. The ingestion plane can degrade (show a queue position) without affecting running inference. The delivery plane can be eventually consistent. But the inference plane must never silently drop a job that consumed GPU time.

Public operating baseline

Midjourney reported approximately 16 million users by mid-2023 and was entirely self-funded, implying massive GPU throughput. Stability AI released Stable Diffusion in August 2022 and SDXL in July 2023, with SDXL generating 1024x1024 images. Runway launched Gen-2 for text-to-video in June 2023. Hugging Face provides serverless GPU inference and dedicated Inference Endpoints supporting A100, H100, and T4 hardware. Replicate's public documentation describes cold-start optimization for SDXL using Cog containers with pre-baked weights. These are public signals that the category is operationally real.

For capacity planning, this answer explicitly assumes a mature consumer service with 500,000 daily active users, 2 million generation requests per day, and a 5x event peak. Unless a number is tied to a citation, it is a stated design assumption.

Key Highlights

  • •GPU inference is a long-running job problem: 4-90 seconds per generation, not a millisecond API call.
  • •Model loading latency (10-30s from network, 2-5s from local NVMe) makes naive autoscaling dangerous.
  • •Four planes: ingestion, orchestration, inference, delivery — each with different consistency and latency needs.
  • •VRAM is the binding constraint, not CPU or network, for diffusion workloads.
  • •Every uncited scale number in this answer is an explicit design assumption, not a company metric.
Lead With GPU Scarcity
State in the first two minutes that the binding constraint is GPU-seconds, not API QPS. This immediately distinguishes a generative AI architecture from a generic web service.
Do Not Design a Synchronous API
A design where the HTTP request blocks until generation completes will timeout under load, waste gateway resources, and fail catastrophically during GPU saturation. The correct pattern is async job submission with polling or webhook delivery.

Section Rescue Kit

Buzzwords to use:

GPU-Second BudgetCold Start Penalty

Safe statements:

  • "I will separate the prompt acceptance path from the GPU execution path because they have fundamentally different latency and failure characteristics."
  • "Before choosing infrastructure, let me define which operations are synchronous (validation, safety check) and which are asynchronous (inference, encoding, delivery)."
Design a Generative Image/Video Service (Stable Diffusion style) - System Design | WinJob | WinJob