Design a Multi-Task Learning System (Unified Model)

Medium45 min
1 / 30
understanding•10 min read

Problem Statement: One Backbone, Many Tasks, Many Failure Modes

Frames the unified multi-task system as a capacity-constrained training and serving platform, not merely one neural network with extra heads.

Problem statement

Design a multi-task learning (MTL) platform that trains, evaluates, releases, and serves one shared model for multiple tasks from different domains: text classification, sentiment analysis, named entity recognition, question answering, and possibly regression tasks such as click-through or dwell-time prediction. The platform must ingest task-specific datasets with different label schemas, combine them into one training schedule, balance task losses so no task dominates, evaluate every task individually and jointly, and serve the correct head for each request while preserving per-task latency and accuracy SLOs.

This is not a single-model architecture question. A strong answer treats the system as four coupled planes. The data plane normalizes heterogeneous task schemas into canonical training records. The training plane schedules task batches, applies adaptive loss weighting or gradient surgery, and manages distributed optimization. The evaluation plane tracks per-task and aggregate metrics, regression gates, and drift. The serving plane routes requests to the shared encoder and the correct task head under strict latency budgets.

Why the problem is distinctive

A single-task pipeline can be tuned independently. A unified system creates cross-task coupling. The classification task may produce large gradients that destabilize the entity-recognition head. The question-answering head may starve the sentiment task because its batches are longer and more expensive. A new task added on Friday may regress an old task by Saturday morning. Therefore the design must separate shared representation learning from task-specific adaptation, and it must make task-level ownership explicit in data, training, evaluation, and serving.

The problem requires a multi-task data pipeline with different label schemas, a shared backbone with separate heads, weighting or scheduling of tasks during training, inference logic that picks the correct head, scalability across many tasks, joint-versus-separate training overhead, reliability when task data updates mid-training, and evaluation on each task individually and combined.

Public operating baseline versus design assumptions

Public evidence shows the category is real. Google's T5 paper cast all text tasks into a text-to-text format and trained one 11-billion-parameter model on many tasks. Meta reported MMoE-style and shared-bottom recommendation models that train multiple objectives in one graph. Apple and Google ship on-device multi-task language models where one encoder serves summarization, classification, and extraction. Hugging Face Transformers reports 1,000,000+ public model repositories and thousands of multi-head checkpoints. These are cited context figures, not requirements for our fictional system.

For capacity planning, this answer explicitly assumes a mature enterprise platform with 40 active tasks, 8,000 combined training examples per second at peak ingestion, 250 million daily inference requests, 25,000 GPU-hours per week of training budget, and a 5× event peak from nightly batch inference and product launches. Unless tied to a citation, every number is a stated design assumption, target, budget, or illustrative threshold.

The four architectural planes

  1. Data plane: task schema registry, canonicalization, validation, sampling budgets, dataset versioning, and mid-training update isolation.
  2. Training plane: shared encoder, task heads, task-aware batching, loss weighting, gradient conflict mitigation, checkpointing, and experiment tracking.
  3. Evaluation plane: per-task metrics, aggregate scorecards, regression gates, ablation baselines, canary evaluation, and drift monitors.
  4. Serving plane: task router, shared encoder service, head registry, batching, quantization, caching, fallback to single-task models, and audit.

A strong interview answer keeps these planes separate. It allows the data plane to add a task without silently changing the training objective, and it allows serving to swap heads without retraining the encoder.

First principles

The core design invariant is: the shared encoder is a capacity budget, and every task is a tenant of that budget. Task data, task loss, task evaluation, and task rollout must be quota-controlled. If one tenant can consume all capacity, the platform degrades to a fragile monolith. If every tenant is isolated, the platform loses the benefit of multi-task learning. The architecture exists to manage this tension explicitly.

Key Highlights

  • •The system is four planes: data, training, evaluation, and serving; the model is only one component.
  • •Cross-task coupling is the central risk: one task's data, loss, or gradient behavior can regress another task.
  • •Assumed mature scale: 40 tasks, 250M inference requests/day, 25K GPU-hours/week, 8K examples/sec peak ingestion.
  • •Public references such as T5, Meta multi-objective ranking, and Hugging Face multi-head ecosystems establish that unified MTL is operationally real.
  • •Treat shared encoder capacity as a budget and each task as a tenant with data, loss, evaluation, and rollout quotas.
Lead With Capacity Tenancy
State early that the shared encoder is a scarce capacity budget and each task is a tenant. This instantly distinguishes a platform design from a toy multi-head neural network.
Do Not Draw One Neural Net and Stop
A diagram with one encoder and many heads but no schema registry, evaluation gates, task routing, or fallback fails a serious interview because it ignores the operational coupling that creates regressions.

Section Rescue Kit

Buzzwords to use:

Negative TransferTask Tenancy

Safe statements:

  • "I will separate model architecture from platform planes because the hard problem is cross-task coupling, not the number of heads."
  • "Before selecting algorithms, let me define which task owns which data, metric, and release gate."
Design a Multi-Task Learning System (Unified Model) - System Design | WinJob | WinJob