How to ground an LLM's answers in your own documents instead of its training data, using the same chunking, embedding, and retrieval pattern that powers production RAG systems in 2026.
Read top to bottom. No signup, no drip โ it's all here.
Retrieval-Augmented Generation means giving a language model access to an external knowledge base at answer time, instead of relying only on what it learned during training. A retriever finds the most relevant chunks of your own documents for a given question, and the model is prompted with the question plus those chunks so it can answer from real, current, source-attributable text rather than a guess.
Per LangChain's own RAG documentation, the pattern is: load documents, split them into chunks, embed the chunks, store them in a vector store, retrieve the most relevant ones for a query, then generate an answer using the question plus the retrieved data. Reach for RAG when your answers need to cite specific, changing, or private documents. Reach for fine-tuning instead when you need the model to change how it behaves or writes, not what facts it knows.
Before anything can be embedded, documents get split into smaller chunks — a whole PDF is too big and too unfocused to embed as one vector. LangChain's docs demonstrate this with RecursiveCharacterTextSplitter, configured with a chunk_size (commonly 1000 characters) and a chunk_overlap (commonly 200 characters) so that context isn't lost at a chunk boundary.
from langchain_text_splitters import RecursiveCharacterTextSplitter
splitter = RecursiveCharacterTextSplitter(
chunk_size=1000,
chunk_overlap=200,
)
chunks = splitter.split_documents(documents)Chunk size is a tuning knob, not a fixed rule. Smaller chunks retrieve more precisely but lose surrounding context; larger chunks keep context but dilute the vector's focus and can pull in irrelevant text alongside the relevant sentence. Start at 1000/200 and adjust based on how your retrieval actually performs on real questions, not on a guess.
An embedding model converts text into a list of numbers (a vector) positioned so that semantically similar text ends up close together in that vector space. Anthropic doesn't ship its own embedding model; its official embeddings guide points to Voyage AI, recommending voyage-4 for balanced general-purpose retrieval, voyage-4-lite for lower latency and cost, and domain-specific models like voyage-law-2 and voyage-code-3.
import voyageai
vo = voyageai.Client()
doc_embds = vo.embed(
documents,
model="voyage-4",
input_type="document",
).embeddingsinput_type to "document" when embedding your knowledge base and "query" when embedding a user's question. The model prepends a different internal instruction for each, and mixing them up measurably hurts retrieval quality.Once chunks are embedded, they need somewhere to live that supports fast similarity search. For prototyping, LangChain's docs use an InMemoryVectorStore that holds everything in memory; for production, the same docs list Chroma, Pinecone, Qdrant, MongoDB, and PGVector as drop-in alternatives behind the same interface.
from langchain_core.vectorstores import InMemoryVectorStore vector_store = InMemoryVectorStore(embedding_model) vector_store.add_documents(chunks) results = vector_store.similarity_search(query, k=4)
Under the hood, similarity search is just comparing the query's vector against every stored vector using a distance metric — commonly cosine similarity or dot product — and returning the closest matches. Voyage's embeddings are normalized to length 1, so per Anthropic's docs, cosine similarity and dot-product similarity give identical rankings; dot product is just cheaper to compute.
Vector similarity search is fast but imprecise — it can rank a chunk that's topically similar above one that actually answers the question. A reranker takes the query plus a shortlist of retrieved documents and re-scores them for relevance using a model built specifically for that comparison, not for embedding. Anthropic's embeddings guide lists rerank-2.5 as Voyage's current recommended reranker for most applications, and rerank-2.5-lite where latency matters more than the last few points of accuracy.
A minimal RAG pipeline is: load your source documents, split them with a text splitter, embed the chunks with input_type="document", store the vectors, embed an incoming question with input_type="query", retrieve the closest chunks (optionally reranked), and pass the question plus those chunks to the model in your prompt.
The most common production failure isn't the model — it's retrieval quietly returning the wrong chunks while the model still writes a confident-sounding answer. Watch for: chunk boundaries that split a sentence containing the actual answer in half, embeddings computed with the wrong input_type, and a vector store that was never re-indexed after the source documents changed. None of these throw an error. They just make the model wrong in a way that looks fine until someone checks the source.
Load one document, split it with RecursiveCharacterTextSplitter, embed the chunks with voyage-4, and print the top match for three test questions.
Run the same document through chunk sizes of 500, 1000, and 2000 characters and compare which one retrieves the right chunk for five known questions.
Retrieve the top 20 chunks by vector similarity, then rerank with rerank-2.5 down to the top 4, and compare answer quality against skipping the rerank step.
Re-implement the starter pipeline against Chroma or PGVector instead of an in-memory store, and measure retrieval latency on a 500+ document corpus.
Build a small script that flags when a source document has changed since it was last embedded and re-indexed, so retrieval never silently serves an outdated chunk.
Sources: LangChain โ Retrieval Augmented Generation (RAG) docs; Anthropic โ Embeddings guide (Voyage AI models, input_type, reranking). Grounded in the official docs and specs above — not invented API details.
This crash course is one of the free courses at Precision AI Academy. No signup, no upsell — just the material.
See all free courses