← BLOG
September 28, 2026

What a RAG Knowledge Base Actually Looks Like in Production

"RAG knowledge base" gets used for both a folder of PDFs pointed at an embedding model and an actual pipeline that stays accurate as your data changes. Here's what separates the two.

"RAG knowledge base" gets used to describe two very different things: a folder of PDFs pointed at an embedding model, and an actual pipeline that stays accurate as the underlying data keeps changing. The first one demos fine in a fifteen-minute call. The second one is what determines whether the system is still trustworthy three months later, after the source documents have already changed twice and nobody remembered to re-ingest them. Here's what actually sits between "we have some documents" and "the answers are still correct in production."

It's a pipeline, not a folder

A knowledge base that only gets built once, at launch, starts going stale the day after launch. What makes it a production system rather than a one-time export is that ingestion, chunking, and storage are a repeatable pipeline, not a manual step someone ran once.

  • Ingestion — pulling from wherever the source of truth actually lives (a CMS, a ticketing system, a database, a shared drive) instead of a manual export someone has to remember to redo.
  • Chunking — splitting documents into pieces small enough to retrieve precisely but large enough to keep their context intact; get this wrong and retrieval returns a sentence with no surrounding meaning to ground it.
  • Metadata — tagging each chunk with what it actually is: source document, section, last-updated date, permission level — so retrieval can filter, not just search everything indiscriminately.
  • Embedding and storage — turning chunks into vectors and storing them somewhere queryable at the latency your product needs, whether that's a managed vector store or something self-hosted.

The parts that don't show up in a demo

A demo only has to work once, on a fixed dataset, in front of an audience that isn't trying to break it. Production has to keep working after the data changes, after two documents disagree with each other, and after someone asks a question the knowledge base genuinely can't answer.

  • Freshness — a re-ingestion trigger tied to the source actually changing, not a nightly cron job that quietly falls behind once document volume grows.
  • Conflict handling — what happens when two source documents say different things about the same topic, and which one the system treats as authoritative.
  • Access control at the chunk level — if different users are supposed to see different subsets of the knowledge base, that has to be enforced at retrieval time, not bolted on afterward as a UI filter.
  • An honest "I don't know" — a system that always returns something, even when nothing relevant exists in the data, is more dangerous than a slower one that says so explicitly.
SEE IT WORKING
Docent — our own RAG platform

Docent is a RAG system we built and run ourselves: grounded answers with citations, an honest refusal instead of a guess when nothing relevant exists, reachable through a dashboard, a public API, or a one-line embeddable widget.

View the project

Keeping it current is the actual hard part

Building the pipeline once is the easy half. The harder, less visible half is noticing when retrieval quality has quietly degraded — because the underlying documents shifted in a way the system wasn't re-indexed for, or because real user questions turned out to be phrased nothing like the queries it was tested on. Without a way to look at actual queries and actual retrieved chunks over time, a knowledge base can look fine in a dashboard and be returning stale or irrelevant context in practice.

This is also where scale changes the shape of the problem. A knowledge base built from a few hundred documents can tolerate a coarse chunking strategy and a simple re-index-everything job. One built from tens of thousands of documents, or a live database updating throughout the day, needs incremental re-indexing, versioning so a bad ingest can be rolled back, and enough monitoring to catch drift before a user does.

What we'd actually recommend

Start with one well-scoped data source and get the full pipeline — ingestion, chunking, metadata, and an honest refusal path — working end to end before expanding to a dozen sources at once. A narrow system that stays accurate beats a broad one that quietly drifts. We scope this after a discovery call, once we've seen what your data actually looks like and how often it changes, since a plan built before that conversation is a guess dressed up as an architecture.

RAG APPLICATIONS

Want to talk through your own build?