Problem Statement: Explanation Compute as a First-Class Service
Frames Model Explanation as a Service as a compute-asymmetric microservice problem, not a thin wrapper around an inference endpoint.
Problem statement
Design a microservice — call it the Explanation Service — that accepts a model reference plus one input instance, runs an interpretability technique (SHAP, LIME, or Integrated Gradients) against that model, and returns feature attributions: a signed contribution per input feature explaining why the model produced its prediction. The service must handle heterogeneous model types (gradient-boosted trees, deep neural networks, opaque third-party API models), track concurrent usage across many models, and durably log requests and partial results for debugging and audit.
This is an emergent MLOps scenario. For years, explanation code lived inside a data scientist's notebook: import shap, explainer = shap.TreeExplainer(model), done. That stops working the moment explanations become a product surface: a fraud analyst needs per-decision reasons inside a case-management UI, a credit officer needs regulator-ready adverse-action factors, an ML platform team needs attribution telemetry for drift detection, and an application backend needs a programmatic API with SLAs. The notebook pattern has no tenancy, no concurrency control, no version pinning, and no SLO. The microservice exists to industrialize what the notebook did.
Why this problem is distinctive
The defining property is compute asymmetry. A single forward pass through an XGBoost model or a small MLP costs microseconds to low milliseconds. An explanation of that same prediction costs orders of magnitude more:
- Tree SHAP is the exception, running in polynomial time O(T·L·D²) for tree ensembles (Lundberg, Erion, Lee and colleagues published the polynomial tree algorithm and its extension to local accuracy in Nature Machine Intelligence in 2020), typically 50–500 ms for a few hundred trees.
- Kernel SHAP approximates Shapley values by evaluating the model on sampled feature coalitions. Exact computation needs 2^M coalition evaluations for M features; practical budgets are 2M + 2048 evaluations (the default background/NSamples settings in the SHAP library). Each evaluation is a model call. For M = 40 features that is roughly 2,100 model calls per explanation.
- LIME fits a local surrogate on perturbed samples — commonly 500 to 5,000 perturbed rows, each scored by the model (Ribeiro, Singh and Guestrin, KDD 2016). For a neural network endpoint this is 5–30 seconds.
- Integrated Gradients accumulates gradients along a straight-line path from a baseline to the input; production systems use 50 to 300 interpolation steps, each a forward-plus-backward pass (Sundararajan, Taly and Yan, ICML 2017). Roughly 1–3 seconds on GPU.
So the same request shape — model ID plus input — maps to workloads that differ by a factor of 10,000 in compute. A service that treats them uniformly will either waste GPUs on tree models or starve LIME jobs behind synchronous traffic. Technique-aware routing is the backbone of this design.
The second hard axis: multi-model concurrency
A platform hosts thousands of registered models but can only hold a fraction in GPU memory or hot CPU memory at once. Explanation requests arrive skewed: ten tenants all explain the same churn model at 9am. The service must decide which models stay resident, when to swap models in and out, how to queue per model without head-of-line blocking across models, and how to cap per-tenant concurrency so one tenant cannot monopolize the pool.
Managed offerings prove the category
AWS documents SageMaker Clarify as the built-in explanation capability of SageMaker, producing SHAP-based attributions in batch transform jobs and, for supported configurations, near-real-time on endpoints. Google Cloud's Vertex AI Explainable AI ships Integrated Gradients, XRAI for images, and Sampled Shapley as attribution methods on deployed models, and its documentation is explicit that approximate methods trade attribution accuracy for latency. TruEra built a standalone AI observability business around scaled SHAP computation and stability monitoring before Snowflake acquired it in May 2023 to fold model explainability into its data platform. The category is real; the interview question asks you to build the engine those products wrap.
Four architectural planes
- Request plane: synchronous and asynchronous API, routing, tenancy, quotas, idempotency.
- Compute plane: technique executors, CPU and GPU worker pools, model session management, admission control.
- Data plane: request ledger, attribution results, durable logs, cache, event backbone.
- Governance plane: model registry integration, version pinning, RBAC, audit, compliance evidence.
A strong answer keeps these planes separate. The compute plane can degrade — slower techniques, queued jobs — without corrupting the request ledger, and the governance plane can block a request without the compute plane ever seeing it.
Key Highlights
- •Explanation compute is 10x to 10,000x a single inference depending on technique — the core design driver.
- •Tree SHAP is polynomial (O(T·L·D²), Lundberg et al., Nature Machine Intelligence 2020); Kernel SHAP exact is 2^M coalitions; LIME is 500–5,000 perturbed scores; IG is 50–300 gradient steps.
- •Managed precedent: SageMaker Clarify (AWS), Vertex AI Explainable AI (Google), TruEra acquired by Snowflake in May 2023.
- •Multi-model concurrency means model residency, swapping and per-model queues are as important as raw worker count.
- •Four planes: request, compute, data, governance. Each degrades independently.
Section Rescue Kit
Buzzwords to use:
Safe statements:
- "I will separate request intake from explanation compute, because their scaling units are completely different."
- "Before choosing infrastructure, let me classify each technique by compute profile, latency budget and model-type compatibility."