Design a Multi-Lingual Machine Translation Service

Hard45 min
1 / 30
understanding•11 min read

Problem Statement: A GPU-Serving Problem Disguised as an API

Frames multi-lingual translation as a model-serving and streaming workload, not a thin proxy over an LLM.

Problem statement

Design a machine translation service that translates text and documents across dozens of languages, in real time and in batch, with domain-specific glossaries for medical, legal, and technical content, streaming output for user input, and a partial feedback loop where human post-edits improve the models. The service backs chat apps, e-commerce listings, support tickets, and document pipelines.

The trap is to draw this as a REST API in front of a model. The actual hard system is a GPU serving platform with a hard latency budget. Translation is autoregressive: output tokens are generated one at a time, so a 40-token sentence from a 600M-parameter multilingual transformer costs roughly 20-40 ms of GPU time per token on an A100-class accelerator. A p50 interactive latency target of under 400 ms means every millisecond of queueing, batching, and network hops must be budgeted, and it means the decode loop, not the HTTP layer, dominates the design.

Why the problem is distinctive

A web service can retry a request. A streaming translation cannot retry the middle of a sentence without the user seeing text rewrite itself. The design therefore separates translation correctness from serving throughput. Correctness is a per-request contract: the right language pair, the right glossary, the right model version, deterministic detokenization. Throughput is a fleet property: continuous batching, replica placement, KV-cache reuse, and autoscaling.

The problem requires multi-lingual model hosting, domain adaptation and glossaries, streaming plus batch modes, and a feedback loop, against non-functional targets for global QPS, low latency, regional reliability, and measurable quality. This answer treats those as architecture requirements, not feature notes.

Public operating baseline versus design assumptions

Public evidence shows the category operates at enormous scale. Google reported in 2016 that Translate handled more than 140 billion words per day, about 1.6 million words per second, and today supports 133 languages. DeepL states it offers 33 languages and that its Pro API supports glossary features across 8 language pairs. Meta's NLLB-200 research covers 200 languages with a single multilingual model and reports a 44% average BLEU improvement over the previous best approach. AWS has published that Amazon Translate has grown by more than 1000% since launch. These are cited company figures and research claims, used as context.

For capacity planning this answer assumes a mature global service with 200M daily active translation events, 10 billion source characters per day, 54 languages with English as a pivot, a 600M-parameter multilingual transformer for interactive traffic, and a 5x peak multiplier. Unless a number is tied to a published source, it is a stated design assumption, target, or budget.

The four architectural planes

  1. Inference plane: GPU replica pools, continuous batching, KV-cache, adaptive quantization, and decode-time token streaming.
  2. Request plane: language detection, script normalization, glossary injection, translation memory lookup, routing, and API/streaming gateways.
  3. Quality plane: BLEU, COMET, BLEURT evaluation pipelines, human evaluation queues, post-edit harvesting, and model canaries.
  4. Data plane: translation memory, glossary stores, parallel corpora, feedback event logs, and the training feature store.

A strong interview answer keeps these planes separate. It allows the request plane to degrade by falling back to a smaller distilled model without touching the quality plane's evaluation gates, and it lets the quality plane improve models without silently changing what is serving live traffic.

Key Highlights

  • •Translation is autoregressive GPU work; the decode loop, not the HTTP layer, sets the latency budget.
  • •A 600M-parameter transformer with 40-token output costs roughly 20-40 ms GPU time per token on A100-class hardware.
  • •Public figures from Google, DeepL, Meta, and Amazon anchor the scale; all other numbers are explicit assumptions.
  • •The architecture has four planes: inference, request, quality, and data.
  • •Fallback to a smaller model is a serving decision; a quality gate prevents that smaller model from ever being promoted silently.
Lead With the Decode Budget
State in the first two minutes that translation latency is dominated by autoregressive decoding on GPUs, so batching, replica placement, and KV-cache are the core design levers. This instantly separates a model-serving design from a generic API design.
Do Not Draw a Thin Proxy
A diagram that is only API gateway, load balancer, and one model box cannot explain how 35K+ requests per second meet a 400 ms p50. The serving fleet and its batching strategy must appear early.

Section Rescue Kit

Buzzwords to use:

Autoregressive DecodingContinuous Batching

Safe statements:

  • "I will separate translation correctness, which is per-request, from serving throughput, which is a fleet property."
  • "Before selecting frameworks, let me define which decisions are made in the request plane, the GPU decode loop, and the quality gate."
Design a Multi-Lingual Machine Translation Service - System Design | WinJob | WinJob