Design a Knowledge Graph Construction from Text (NLP)

Hard45 min
1 / 30
understanding•11 min read

Problem Statement: Turning Noisy Text Into a Queryable Graph

Frames the pipeline as a probabilistic-to-deterministic bridge with four planes: ingestion, extraction, assembly, and serving.

Problem statement

Design a pipeline that reads large corpora of text (news articles, internal documents, reports), runs NLP to identify entities (people, organizations, places, products, financial instruments) and the relationships between them, and assembles the results into an evolving knowledge graph. The graph must handle partial entity merges, references to the same real-world object across documents, versioning when source text changes, storage in a specialized graph database, and a query interface for search and analysis.

The defining difficulty is that the two halves of the system have opposite contracts. The NLP half is probabilistic: every extracted mention and relation carries a confidence score, and precision and recall move with model versions and corpus drift. The graph half must present a deterministic, queryable surface: an analyst asking "which companies does this person sit on the board of" needs one authoritative answer with sources, not a distribution. The architecture is therefore a controlled bridge from noisy evidence to stable facts, and every subsystem in this answer exists to make that bridge safe: provenance so facts can be traced to text, confidence so weak evidence can be filtered, entity resolution so duplicates collapse, and versioning so truth can be reconstructed at any point in time.

Why the problem is distinctive

A search engine can return a ranked list and be done. A knowledge graph cannot. If the extractor links "A. Smith" in one document and "Adam Smith" in another to different nodes when they are the same person, every downstream consumer inherits the split forever. If it over-merges two different people, the graph asserts a falsehood that no query can detect. Entity resolution errors compound multiplicatively through edges. Likewise, when a source document is corrected, the graph must know which facts depended on the old text and retire them without a full rebuild. This is why the brief explicitly calls out partial entity merges, cross-document references, and versioning: they are the failure modes, not edge cases.

Public operating baseline versus design assumptions

Public evidence establishes that this category is real and large. Google reported at the May 2012 launch of its Knowledge Graph roughly 500 million entities and more than 50 billion facts, under the "things, not strings" framing published by Ramanathan Guha. Wikidata has passed 100 million community-curated items. Microsoft maintained the Microsoft Academic Graph, on the order of 250 million papers and billions of citation edges, as a public dataset before retiring the service at the end of 2021. Diffbot claims an automatically assembled web-scale graph with roughly 10 billion entity profiles built from web crawls. These are cited public figures for context, not requirements for our fictional system.

For capacity planning, this answer assumes a mature enterprise graph: 50 million documents under management, 500 thousand document events per day (new plus updated), a canonical graph of 400 million entities and 8 billion edges, and 2,000 interactive queries per second. Unless a number is tied to a citation above, it is a stated design assumption, target, or budget.

The four architectural planes

  1. Ingestion plane: corpus onboarding, document normalization, content deduplication, document version tracking, and re-extraction triggers.
  2. Extraction plane: NER, coreference, relation extraction, and confidence calibration, executed on GPUs with model version pinning.
  3. Assembly plane: entity resolution, cluster management, edge assembly, conflict resolution, provenance attachment, and graph versioning.
  4. Serving plane: graph storage, indexes, vector search for entity linking, query APIs, snapshots, and analytics projections.

A strong interview answer keeps these planes separate. Extraction may be re-run freely because its outputs are immutable evidence. Assembly is the only plane allowed to mutate canonical identity. Serving reads only published, versioned state. That separation is what makes reprocessing cheap, merges auditable, and regressions reversible.

Key Highlights

  • •The core tension: probabilistic extraction must feed a deterministic, queryable graph surface.
  • •Entity resolution errors compound: one bad merge or split corrupts every downstream edge.
  • •Public anchors: Google reported ~500M entities and ~50B facts at 2012 Knowledge Graph launch; Wikidata passed 100M items.
  • •Four planes: ingestion, extraction, assembly, serving. Only assembly mutates canonical identity.
  • •Every uncited scale number in this answer is an explicit design assumption, not a company fact.
Lead With the Probabilistic-to-Deterministic Bridge
State in the first two minutes that extraction is probabilistic evidence while the graph must be an authoritative surface, and that provenance plus confidence is what makes the bridge safe. This instantly separates an NLP-systems answer from a generic ETL answer.
Do Not Draw a Batch ETL Toy
A design where an extractor writes directly into the graph database, with no mention store, no confidence, no provenance, and no versioning, cannot answer 'why does the graph say this' or survive a model regression. It will fail a serious interview.

Section Rescue Kit

Buzzwords to use:

Things, Not StringsProvenance-Anchored Fact

Safe statements:

  • "I will separate evidence from assertion: extractors emit immutable mention events, and only the assembly plane writes canonical graph state."
  • "Before choosing databases, let me define which plane owns identity, which owns evidence, and which owns publication."
Design a Knowledge Graph Construction from Text (NLP) - System Design | WinJob | WinJob