Design a Fraud Detection Microservice

Hard45 min
1 / 30
understanding•8 min read

Problem Statement: Real-Time Risk Screening at Checkout

Frames fraud detection as a latency-bound decisioning service, not a batch analytics job.

Problem statement

Design a fraud detection microservice that screens every e-commerce order at checkout. The order pipeline sends a transaction context: shopper identity signals, device fingerprint, IP and geolocation, cart contents, payment token, merchant and shipping details. The service must assemble risk features, evaluate rules, run a machine-learning score, and return a decision - approve, decline, or send to manual review - within a strict latency budget so checkout is not blocked.

This is a decisioning system, not a reporting system. A declined order is a lost customer if it was legitimate, and an approved fraudulent order becomes a chargeback, a fee, and inventory loss. The service therefore owns three distinct loops:

  1. Synchronous decision loop: features, rules, model score, decision, all inside the checkout latency budget.
  2. Asynchronous investigation loop: device-ring and link analysis, review cases, analyst overrides.
  3. Learning loop: labels from chargebacks and reviews, training data assembly, model retraining and governed rollout.

Why the problem is distinctive

Checkout cannot wait. If the fraud service is slow or down, the merchant either blocks all orders (unacceptable) or approves everything (also unacceptable). The design must define explicit degraded modes with a chosen risk posture. Second, the adversary adapts: a static ruleset decays within weeks, so the learning loop and model refresh cadence are architecture requirements, not ML team details. Third, decisions must be explainable and auditable: an analyst reviewing a case, a regulator asking about decline reasons, and a model debugger replaying an incident all need the exact feature snapshot, rule version, and model version that produced the decision.

Public operating signals show the category is real and scaled. Visa publicly describes analyzing hundreds of risk attributes in roughly one millisecond per authorization at network-scale throughput. Stripe publicly describes Radar as an ML system trained on data from millions of companies and exposing a rules language for merchants. Forter publicly markets instant approve/decline decisions backed by a chargeback guarantee. These are company-reported figures for context, not requirements for our design.

Design assumptions

For capacity math this answer assumes a mid-size marketplace: 20,000,000 checkout attempts per day, a 5x event peak, fraud scoring budget of 80 ms p99 inside a 150 ms checkout budget, and roughly 0.8% of orders routed to manual review. Every uncited number is an explicit assumption, target, or budget - not a claim about any company's private architecture.

The four architectural planes

  1. Decision plane: synchronous scoring - feature assembly, rules engine, model inference, decision policy.
  2. Signals plane: device fingerprinting, IP intelligence, velocity counters, entity profiles, feature store.
  3. Review operations plane: case queues, analyst tooling, overrides, evidence.
  4. Learning plane: label ingestion, training data, model evaluation, champion/challenger rollout, drift monitoring.

A strong answer keeps these planes separate. The review plane can be slow without hurting checkout. The learning plane can lag without weakening the current decision path. The decision plane must stay fast and correct even when the other three degrade.

Key Highlights

  • •Fraud detection at checkout is a latency-bound decisioning service, not batch analytics.
  • •Three loops: synchronous decision, asynchronous investigation, and governed learning.
  • •Every decision must replay exactly: feature snapshot, rule version, model version.
  • •Company-public scale signals are context; all capacity numbers here are stated assumptions.
  • •Four planes: decision, signals, review operations, learning.
Lead With the Latency Budget
State in the first two minutes that the decision must land inside the checkout latency budget, and that a degraded mode must be defined for when scoring is unavailable. This immediately separates a decisioning system from an analytics pipeline.
Do Not Design a Nightly Batch Job
Scoring orders hours after checkout cannot prevent fraud or chargebacks; the order has already shipped. Batch belongs only in the learning and investigation planes.

Section Rescue Kit

Buzzwords to use:

Decision BudgetRisk Posture

Safe statements:

  • "I will separate the synchronous decision loop from investigation and learning before drawing services."
  • "Let me state the latency budget and degraded posture first, because they constrain every later choice."
Design a Fraud Detection Microservice - System Design | WinJob | WinJob