The complete pipeline
A production flow begins before the user asks a question: collect approved sources, normalize them, split them into meaningful units, attach metadata, create embeddings, and store both text and vectors. At query time, retrieve candidates, optionally filter or rerank them, build a bounded context, generate an answer, and return citations.
Chunking is an information-design decision
A fixed character count is a starting point, not a strategy. Preserve headings, lists, tables, and semantic boundaries. A chunk should be specific enough to retrieve precisely and complete enough to answer without inventing the missing half.
Retrieval and generation fail differently
If the right evidence is absent from the retrieved set, rewriting the final prompt cannot repair retrieval. Measure whether relevant evidence appears in the top results before judging the generated answer. Then measure faithfulness: does the answer stay within that evidence?
Abstention is a feature
The system needs permission to say that the available sources do not support an answer. Define minimum relevance, require citations for factual claims, and provide an escalation path when evidence conflicts or is missing.
Production checklist
Version documents and embeddings, preserve source identifiers, enforce access control before retrieval, log which chunks supported each response, remove sensitive data, test prompt injection inside documents, and monitor cost and latency as the collection grows.