Free Crash Course ยท RAG

Learn RAG: Retrieval-Augmented Generation From Scratch

How to ground an LLM's answers in your own documents instead of its training data, using the same chunking, embedding, and retrieval pattern that powers production RAG systems in 2026.

Free crash course
Everything on this page
No signup
Self-paced
6
Lessons
5
Projects
3
Real API Docs
$0
Forever Free

What this course covers.

6 lessons, everything on this page.

Read top to bottom. No signup, no drip โ€” it's all here.

1

1. What RAG is, and when you actually need it

Retrieval-Augmented Generation means giving a language model access to an external knowledge base at answer time, instead of relying only on what it learned during training. A retriever finds the most relevant chunks of your own documents for a given question, and the model is prompted with the question plus those chunks so it can answer from real, current, source-attributable text rather than a guess.

Per LangChain's own RAG documentation, the pattern is: load documents, split them into chunks, embed the chunks, store them in a vector store, retrieve the most relevant ones for a query, then generate an answer using the question plus the retrieved data. Reach for RAG when your answers need to cite specific, changing, or private documents. Reach for fine-tuning instead when you need the model to change how it behaves or writes, not what facts it knows.

Rule of thumb: if the honest test of a good answer is "did it get the facts from the right document," that's RAG. If the test is "did it write in the right style or follow the right format," that's fine-tuning.
2

2. Chunking: splitting documents so they retrieve well

Before anything can be embedded, documents get split into smaller chunks — a whole PDF is too big and too unfocused to embed as one vector. LangChain's docs demonstrate this with RecursiveCharacterTextSplitter, configured with a chunk_size (commonly 1000 characters) and a chunk_overlap (commonly 200 characters) so that context isn't lost at a chunk boundary.

from langchain_text_splitters import RecursiveCharacterTextSplitter

splitter = RecursiveCharacterTextSplitter(
    chunk_size=1000,
    chunk_overlap=200,
)
chunks = splitter.split_documents(documents)

Chunk size is a tuning knob, not a fixed rule. Smaller chunks retrieve more precisely but lose surrounding context; larger chunks keep context but dilute the vector's focus and can pull in irrelevant text alongside the relevant sentence. Start at 1000/200 and adjust based on how your retrieval actually performs on real questions, not on a guess.

3

3. Embeddings: turning text into vectors

An embedding model converts text into a list of numbers (a vector) positioned so that semantically similar text ends up close together in that vector space. Anthropic doesn't ship its own embedding model; its official embeddings guide points to Voyage AI, recommending voyage-4 for balanced general-purpose retrieval, voyage-4-lite for lower latency and cost, and domain-specific models like voyage-law-2 and voyage-code-3.

import voyageai

vo = voyageai.Client()
doc_embds = vo.embed(
    documents,
    model="voyage-4",
    input_type="document",
).embeddings
Gotcha: Anthropic's docs stress always setting input_type to "document" when embedding your knowledge base and "query" when embedding a user's question. The model prepends a different internal instruction for each, and mixing them up measurably hurts retrieval quality.
4

4. Storing and searching vectors

Once chunks are embedded, they need somewhere to live that supports fast similarity search. For prototyping, LangChain's docs use an InMemoryVectorStore that holds everything in memory; for production, the same docs list Chroma, Pinecone, Qdrant, MongoDB, and PGVector as drop-in alternatives behind the same interface.

from langchain_core.vectorstores import InMemoryVectorStore

vector_store = InMemoryVectorStore(embedding_model)
vector_store.add_documents(chunks)

results = vector_store.similarity_search(query, k=4)

Under the hood, similarity search is just comparing the query's vector against every stored vector using a distance metric — commonly cosine similarity or dot product — and returning the closest matches. Voyage's embeddings are normalized to length 1, so per Anthropic's docs, cosine similarity and dot-product similarity give identical rankings; dot product is just cheaper to compute.

5

5. Reranking: fixing what plain vector search gets wrong

Vector similarity search is fast but imprecise — it can rank a chunk that's topically similar above one that actually answers the question. A reranker takes the query plus a shortlist of retrieved documents and re-scores them for relevance using a model built specifically for that comparison, not for embedding. Anthropic's embeddings guide lists rerank-2.5 as Voyage's current recommended reranker for most applications, and rerank-2.5-lite where latency matters more than the last few points of accuracy.

Practical pattern: retrieve a wider set first (say, the top 20 chunks by vector similarity), then rerank down to the 3-5 you actually pass to the model. This two-stage retrieve-then-rerank approach consistently beats using either step alone.
6

6. Putting it together, and where RAG breaks in production

A minimal RAG pipeline is: load your source documents, split them with a text splitter, embed the chunks with input_type="document", store the vectors, embed an incoming question with input_type="query", retrieve the closest chunks (optionally reranked), and pass the question plus those chunks to the model in your prompt.

The most common production failure isn't the model — it's retrieval quietly returning the wrong chunks while the model still writes a confident-sounding answer. Watch for: chunk boundaries that split a sentence containing the actual answer in half, embeddings computed with the wrong input_type, and a vector store that was never re-indexed after the source documents changed. None of these throw an error. They just make the model wrong in a way that looks fine until someone checks the source.

Build these to make it stick.

Starter

Chunk and embed a single PDF

Load one document, split it with RecursiveCharacterTextSplitter, embed the chunks with voyage-4, and print the top match for three test questions.

Starter

Compare chunk sizes on the same document

Run the same document through chunk sizes of 500, 1000, and 2000 characters and compare which one retrieves the right chunk for five known questions.

Intermediate

Build a two-stage retrieve-then-rerank pipeline

Retrieve the top 20 chunks by vector similarity, then rerank with rerank-2.5 down to the top 4, and compare answer quality against skipping the rerank step.

Intermediate

Swap InMemoryVectorStore for a real vector database

Re-implement the starter pipeline against Chroma or PGVector instead of an in-memory store, and measure retrieval latency on a 500+ document corpus.

Advanced

Detect stale retrieval after a document update

Build a small script that flags when a source document has changed since it was last embedded and re-indexed, so retrieval never silently serves an outdated chunk.

Sources: LangChain โ€” Retrieval Augmented Generation (RAG) docs; Anthropic โ€” Embeddings guide (Voyage AI models, input_type, reranking). Grounded in the official docs and specs above — not invented API details.

Where to head next.

Free, and staying that way.

This crash course is one of the free courses at Precision AI Academy. No signup, no upsell — just the material.

See all free courses