Design a Document QA System with Hybrid Retrieval (Keyword + Vector)

Medium45 min
1 / 31
understanding•10 min read

Problem Statement: Grounded Answers Over a Large Document Corpus

Frames the product as a retrieval-grounded generation pipeline where retrieval quality, not model size, determines answer quality.

Problem statement

Design a document question-answering system that answers natural-language questions over a large, continuously updated document corpus - product documentation, wiki pages, policies, tickets, PDFs, and markdown - by combining classic keyword retrieval with dense vector retrieval, fusing the two result lists, reranking the union, and feeding the top passages to a large language model that writes a grounded, cited answer. The system must handle concurrent queries, partial document updates, multi-turn sessions, user feedback, and UI snippet highlighting.

This is not a chatbot prompt-engineering exercise. The dominant failure mode in production QA is retrieval failure: the model writes fluently about passages it never saw, or the right paragraph was ranked 40th and never reached the context window. Anthropic's published Contextual Retrieval analysis reports that adding contextual BM25 to contextual embeddings plus a reranker reduced top-20 retrieval failures by 67% compared with embeddings alone, which is the strongest public signal that hybrid retrieval is not optional polish but the core architecture. Microsoft's Azure AI Search team published the same direction of result: hybrid retrieval plus a semantic reranker outperformed pure vector search on their benchmarks.

Why hybrid

Keyword retrieval (BM25 over an inverted index) is exact, cheap, deterministic, and unbeatable for rare terms: error codes, SKU identifiers, API names, proper nouns, acronyms. Dense vector retrieval survives paraphrase, synonyms, and cross-lingual rewording but misses exact tokens and can misrank rare identifiers. The two failure modes are nearly orthogonal, which is precisely why rank fusion works. A question like 'why does Error E-4012 appear when syncing a workspace' needs BM25 to find E-4012 at all, and vectors to find the paragraph that explains it in different words.

The pipeline in one breath

Ingest documents, parse and normalize them, split them into passages with preserved lineage, embed each passage, and write both an inverted index entry and a vector into serving stores. At query time: rewrite the question if needed, run BM25 and approximate nearest-neighbor search in parallel, fuse the two ranked lists with Reciprocal Rank Fusion, cross-encoder rerank the fused candidates, pack the survivors into the model context under a token budget, generate a streaming answer with per-sentence citations, and capture feedback.

What makes it hard

  1. Retrieval correctness under freshness pressure: a document edited five minutes ago must not contradict tonight's answers, so indexing is a streaming pipeline with per-document version pinning, not a nightly batch.
  2. Latency composition: retrieval, rerank, and generation run sequentially inside one user-perceived budget of a few seconds, and each stage has its own tail.
  3. Concurrency across a multi-tier pipeline: one user query fans out to two index shards, one reranker GPU call, and one LLM stream; backpressure anywhere must degrade gracefully to keyword-only rather than fail.
  4. Grounding and trust: every claim maps back to a passage ID so the UI can highlight snippets and an evaluator can audit hallucination.
  5. Cost: an embedding model, a reranker, and a generator all cost money per query, so caching and admission control are architecture, not afterthoughts.

The four planes

  1. Ingestion plane: parsing, chunking, enrichment, embedding, dual-index writing, version pinning.
  2. Retrieval plane: query understanding, parallel keyword and vector recall, fusion, reranking.
  3. Generation plane: context assembly, grounded decoding, citations, streaming, refusal.
  4. Learning plane: feedback capture, retrieval and answer evaluation, ranking supervision, index quality monitoring.

A strong answer keeps these planes separate: the generation plane may degrade to a shorter answer, the retrieval plane may degrade to keyword-only, but the ingestion plane's version contract never silently weakens, because stale passages poison every downstream answer.

Key Highlights

  • •Retrieval failure, not generation failure, is the dominant production failure mode; Anthropic reports a 67% reduction in top-20 retrieval failures when contextual BM25 plus reranking is added to embeddings.
  • •BM25 and vector search have near-orthogonal failure modes: exact rare tokens versus paraphrase, which is the entire argument for hybrid recall.
  • •One query fans out across two index systems, a GPU reranker, and an LLM stream inside a single multi-second latency budget.
  • •Every generated sentence must carry passage-level lineage so the UI can highlight snippets and evaluators can audit grounding.
  • •Indexing is a streaming, version-pinned pipeline; a stale chunk is a correctness bug, not a performance issue.
Lead With Retrieval, Not the LLM
State in the first two minutes that answer quality is bounded by retrieval quality, and that hybrid recall plus reranking is the mechanism. This instantly separates a RAG architecture answer from a prompt-engineering answer.
Do Not Draw One Magic Search Box
A design with a single unspecified search service hides the entire interesting problem: two indexes with different consistency, latency, and scaling behavior, plus a fusion contract between them.

Section Rescue Kit

Buzzwords to use:

Reciprocal Rank FusionGrounded Generation

Safe statements:

  • "I will separate answer generation from evidence retrieval, because the two fail differently and scale differently."
  • "Before choosing any model, let me define what the retrieval layer must guarantee for freshness and recall."
Design a Document QA System with Hybrid Retrieval (Keyword + Vector) - System Design | WinJob | WinJob