← All Articles
RAGLLM EngineeringAI Infrastructure

Building Production-Grade RAG Systems: The Architecture Decisions That Actually Matter

Christian Chukwuka··5 min read
Building Production-Grade RAG Systems: The Architecture Decisions That Actually Matter
TL;DR

Production RAG fails on chunking, not prompting: fixed-size windows split tables and clauses, pure vector search misses exact IDs, and most teams ship with no retrieval-level eval. Fix: semantic chunking, hybrid BM25+vector+reranking, and a separate eval harness for retrieval vs. generation quality.

Every RAG tutorial follows the same shape: load a PDF, chunk it, embed it, drop it in a vector store, retrieve top-k, stuff it into a prompt. It works great in a demo with twelve pages of clean text. It falls apart within a month of real usage — not because the model is weak, but because the retrieval pipeline around it was never designed as a system.

Chunking is not a solved problem

Fixed-size chunking with a sliding window is the default in almost every framework, and it is the first thing that breaks. A 500-token chunk boundary that splits a table in half, or separates a clause from the condition that governs it, produces retrieval that looks correct and is quietly wrong.

  • Semantic chunking (splitting on structural boundaries — headings, list items, table rows) outperforms fixed windows on anything structured: contracts, invoices, technical docs.
  • Overlap should be a function of document type, not a global constant. Narrative text tolerates 10-15% overlap; structured data often needs none, and instead needs metadata (page, section, parent heading) attached to every chunk.
  • Chunk size is a latency and cost decision as much as a retrieval-quality one — larger chunks mean fewer retrieved items fit in context, and more tokens billed per query.

Why pure vector search fails in production

Dense retrieval is excellent at semantic similarity and bad at exact match. Ask it to find "invoice INV-2291" and it will happily return five semantically related invoices that are not INV-2291, because embeddings do not privilege exact tokens. Any production RAG system that handles IDs, SKUs, ticket numbers, or names needs a lexical layer, not just a vector one.

Hybrid retrieval: BM25 + vectors + reranking

The pattern that actually holds up is hybrid search — combine a lexical scorer like BM25 with dense vector similarity, merge the candidate sets, and pass the merged set through a cross-encoder reranker before it ever reaches the LLM. The reranker is doing the expensive, accurate comparison on a small candidate set instead of the whole corpus, which keeps latency bounded.

retrieval_pipeline.py
candidates = bm25_search(query, k=50) + vector_search(query, k=50)
candidates = dedupe(candidates)
ranked = reranker.score(query, candidates)  # cross-encoder
top_k = ranked[:8]
context = build_context(top_k, max_tokens=3000)

Evaluation is the part everyone skips

Teams ship RAG systems with no regression test suite, then get surprised when a prompt tweak silently degrades retrieval quality for an entire category of queries. A minimal eval harness needs retrieval-level metrics (precision@k, recall@k against a labeled query set) separate from generation-level metrics (an LLM-as-judge grading faithfulness and completeness). Conflating the two hides which layer actually broke.

Cost and latency are architecture decisions, not afterthoughts

Every additional reranking pass, every larger context window, every synchronous embedding call on the request path is a latency and a cost line item. The teams that get this right treat retrieval depth as a tunable parameter validated against the eval harness, not a constant picked once and forgotten.

Stale source documents are a silent failure mode

A retrieval pipeline that was accurate at launch degrades as the underlying documents change — a policy gets updated, a product spec changes, a page gets deleted — and unless re-indexing is an explicit, monitored process, the system keeps confidently retrieving and citing the old version. This failure is worse than a retrieval miss, because it doesn't look like a failure: the system returns a well-formed, plausible, and wrong answer sourced from a document that's no longer current. Re-indexing on a schedule tied to how often the source actually changes — not a fixed interval picked once — and tracking document version alongside each chunk is what catches this before a user does.

Retrieval quality needs its own eval, not just an end-to-end score

An end-to-end eval that only grades the final generated answer can't tell you whether a bad answer came from bad retrieval or a bad generation given good context — and those have completely different fixes. Separating retrieval-level metrics (precision@k, recall@k against a labeled query set) from generation-level metrics (faithfulness to the retrieved context, completeness) is what turns "the RAG system got worse" into "retrieval degraded on this category of query" or "generation started ignoring context it was given" — genuinely different bugs.

For the storage layer underneath most of this, see Postgres as a vector store. For building the labeled query sets this kind of retrieval eval actually needs, see setting up golden datasets for LLM regression testing, and for the production monitoring side, LLM observability in production.

Frequently asked questions

How often should a RAG index be refreshed?

It should be tied to how frequently the underlying source documents actually change, not a fixed interval — a knowledge base that updates daily needs a very different re-indexing cadence than a set of contracts that rarely change. The more important practice is tracking document version alongside each chunk so staleness is at least detectable, even if refresh isn't instant.

Why does my RAG system return confident but outdated answers?

This is almost always a stale-index problem, not a generation problem — the retrieval layer is returning chunks from a document version that's no longer current, and the model is faithfully summarizing what it was given. Fixing the generation prompt won't help; the fix is re-indexing on a cadence that matches how often the source actually changes.

The takeaway

RAG that works in production is built by people who treat it as an information retrieval system with an LLM at the end, not a prompt-engineering exercise with a database attached. Chunking strategy, hybrid retrieval, reranking, and continuous evaluation are the actual engineering — the prompt is the easy 10%.

Christian Chukwuka
Christian Chukwuka
Founder & AI Systems Engineer

Have a similar challenge?

Book a free 30-minute architecture call and we'll tell you honestly whether and how we can help.

Book Free Discovery Call →