Design a Chatbot for Customer Support with LLM & Knowledge Base

Hard45 min
1 / 30
understanding•10 min read

Problem Statement: A Grounded, Fresh, and Accountable Support Agent

Frames the chatbot as a retrieval-augmented generation system with freshness, groundedness, and escalation contracts, not a wrapped LLM call.

Problem statement

Design a customer-support chatbot that converses naturally with an LLM while answering only from a company knowledge base that changes every day. The system must ingest articles in near real time, retrieve the right passages per turn, assemble a bounded prompt, generate a grounded streamed answer with citations, detect its own uncertainty, and hand off to a human with full context when confidence or policy demands it.

The defining difficulty is not 'call an LLM'. It is that three clocks run at different speeds: the conversation clock (sub-second token streaming), the knowledge clock (articles published, edited, and retracted minutes ago), and the trust clock (every answer must be defensible, citable, and auditable after the fact). A design that treats the knowledge base as a static blob re-embedded nightly fails the brief; a design that re-embeds everything on every edit fails on cost; a design that lets the model answer from parametric memory fails on correctness.

Why this problem is distinctive

A search engine tolerates a stale index for hours. A support bot cannot: a pricing change, a security incident page, or a retracted refund policy must stop being quoted within minutes, and must never have been quoted after retraction. Meanwhile the generator is stochastic: the same retrieval can produce an unsupported claim. Therefore the architecture separates four planes: the conversation plane (sessions, turns, streaming), the knowledge plane (CMS, chunking, embeddings, index versions), the inference plane (model router, guardrails, semantic cache), and the assurance plane (evals, traces, citations, escalation audit). Each plane has its own consistency model, its own failure ladder, and its own release gate.

Public operating baseline versus design assumptions

Public evidence shows the category operates at real scale. Klarna reported that its OpenAI-powered assistant handled 2.3 million conversations in its first month, roughly two thirds of customer service chats, performing the equivalent work of about 700 full-time agents and projecting around 40 million USD in profit improvement. Intercom reports that its Fin agent typically resolves around half of support conversations for customers, with top accounts materially higher. Salesforce shipped Agentforce with the Atlas Reasoning Engine and the Einstein Trust Layer, which contractually avoids retaining customer prompts for model training. These are cited company-reported figures, not requirements for our fictional system.

For capacity planning this answer explicitly assumes a mature support organization: 2,000,000 conversations per month, 8 user turns per conversation, 250,000 knowledge articles across 12 locales, 5,000 article updates per day, 5,000 concurrent sessions at peak, and an 8x peak multiplier on turn rate. Unless a number is tied to a citation, it is a stated design assumption, target, or budget.

The four architectural planes

  1. Conversation plane: channel gateway, session state machine, turn orchestration, streaming delivery, escalation handoff.
  2. Knowledge plane: CMS webhooks, parsing, chunking, embedding, versioned index aliases, freshness SLOs.
  3. Inference plane: hybrid retrieval, reranking, prompt assembly, semantic cache, model router, input and output guardrails.
  4. Assurance plane: citation binding, faithfulness evals, offline and online evaluation, trace retention, incident review, release gates.

A strong interview answer keeps these planes separate, states which decisions are synchronous on the turn path and which are asynchronous, and names the invariant: the bot may only assert what the retrieved, version-pinned evidence supports, and must escalate rather than guess.

Key Highlights

  • •Three clocks dominate: sub-second conversation streaming, minute-scale knowledge freshness, and post-hoc auditability of every claim.
  • •Klarna reported 2.3M conversations in month one, about 700 FTE equivalence; Intercom reports Fin resolving roughly half of conversations; both are company-reported figures.
  • •Design assumptions: 2M conversations/month, 8 turns each, 250K articles, 5K updates/day, 5K concurrent sessions, 8x peak.
  • •Four planes: conversation, knowledge, inference, assurance; each has its own consistency model and failure ladder.
  • •The invariant: assert only what version-pinned retrieved evidence supports; escalate instead of guessing.
Lead With Groundedness, Not Fluency
State in the first two minutes that the bot may only assert what retrieved, version-pinned evidence supports, and that abstention plus escalation is a designed outcome, not a failure. This separates a production RAG design from a demo wrapper.
Do Not Assume a Nightly Re-Embed
A nightly full re-index means a retracted refund policy keeps being quoted for up to 24 hours. The brief explicitly demands real-time ingestion; design delta upserts plus alias swaps from the start.

Section Rescue Kit

Buzzwords to use:

Retrieval-Augmented GenerationGroundedness Contract

Safe statements:

  • "I will separate conversation latency, knowledge freshness, and answer accountability before choosing any component."
  • "Before drawing services, let me state what the bot is allowed to assert and what it must escalate instead."
Design a Chatbot for Customer Support with LLM & Knowledge Base - System Design | WinJob | WinJob