Problem Statement: A Clinical AI Platform, Not a Model
Frames the system as a regulated, hospital-integrated diagnosis support platform where the transformer is one component among ingestion, governance, serving, and validation planes.
Problem statement
Design a platform that trains, validates, deploys, and monitors a Vision Transformer (ViT) that reads labeled medical images — chest X-rays and MRI slices — and supports clinicians by flagging suspected conditions, attaching evidence regions, and prioritizing urgent studies. The system ingests studies from hospital PACS over DICOM, de-identifies them under HIPAA, manages noisy labels derived from radiology reports, trains and evaluates transformer models against locked multi-site test sets, serves inferences inside clinical latency budgets, and returns results as DICOM Structured Reports, segmentation objects, and FHIR DiagnosticReport entries the radiologist workflow can consume.
This is not a Kaggle classification exercise. A chest radiograph from a computed radiography detector is typically a 3000 x 3000 pixel, 12-to-16-bit image of 10 to 50 MB, while a ViT consumes fixed 224, 384, or 512 pixel patches. The first hard problem is therefore resolution bridging: tiling, hierarchical encoders, or global-local fusion — chosen with an evidence budget, not aesthetics. The second is that ground truth is uncertain: labels are mined from free-text reports, disagree between readers, and carry implied uncertainty the way CheXpert encoded U-Ones and U-Zeros. The third is clinical integration: the model output must reach a radiologist inside their existing viewer with an explainable overlay, and their agreement or override must flow back as supervised signal. The fourth is regulatory: a locked, validated algorithm with documented performance across sites, sexes, age bands, and device vendors — because FDA-cleared AI behaves differently from a continuously redeployed web model.
Public evidence baseline
The category is real and published. The NIH ChestX-ray14 release contains 112,120 frontal-view chest X-rays from 30,805 patients. Stanford's CheXpert contains 224,316 chest radiographs from 65,240 patients and explicitly models label uncertainty. MIMIC-CXR provides 377,110 images across 227,827 studies under credentialed access. Google Health's Nature 2020 breast-cancer screening study evaluated 25,856 UK and 3,097 US screening mammograms and reported false-positive reductions of 5.7% and 1.2% and false-negative reductions of 9.4% and 2.7% versus single radiologists. The FDA's public list of authorized AI/ML-enabled medical devices exceeded 900 entries by late 2023, with radiology the dominant specialty. These are cited context figures; every uncited scale number in this answer is an explicit design assumption.
The five architectural planes
- Clinical integration plane: DICOM and HL7 connectivity, worklist awareness, result delivery, viewer overlays.
- Data governance plane: de-identification, lineage, consent, quarantine, retention.
- Training plane: dataset versioning, distributed fine-tuning, evaluation harness, experiment tracking.
- Serving plane: model registry, in-hospital or regional inference, DICOM SR/SEG authoring, priority routing.
- Validation and monitoring plane: locked test sets, drift detection, override feedback, regulatory evidence.
A strong answer keeps these planes separate. A model improvement must not silently change the validated serving contract, and a PACS outage must not corrupt the training dataset or the audit trail.
Key Highlights
- •The deliverable is a regulated clinical decision-support platform; the ViT is one component inside five planes.
- •Chest radiographs arrive as 3000 x 3000, 10-50 MB DICOM objects while ViT consumes 224-512 px patches: resolution bridging is the core ML engineering problem.
- •Ground truth is uncertain and report-derived; the label model is a first-class subsystem, not preprocessing.
- •NIH ChestX-ray14 (112,120 images), CheXpert (224,316 images), and MIMIC-CXR (377,110 images) anchor public dataset scale.
- •Every uncited number in this answer is an explicit design assumption, target, or budget.
Section Rescue Kit
Buzzwords to use:
Safe statements:
- "I will separate model quality from platform correctness: a high-AUC model that cannot reach a PACS or survive an audit is not a product."
- "Before choosing architectures, let me define which decisions belong to the hospital, the ML platform, and the regulator."