This lesson on Retrieval-Augmented Generation (RAG) Foundations is hands-on and example-driven. You will build and deploy an end-to-end Retrieval-Augmented Generation (RAG) pipeline in Python that bridges static LLM knowledge cutoffs with proprietary external data. You will ingest unstructured text, generate dense embeddings across compatible vector dimensions, and inject relevant context into generation prompts to eliminate hallucinations.
What You'll Be Able To Do
- Identify LLM knowledge boundary limitations and hallucination risks caused by static pretraining cutoffs.
- Generate dense vector representations of text using local Ollama and remote embedding models across exact dimensional specifications.
- Execute semantic similarity searches against vector stores like Chroma DB to retrieve relevant context chunks.
- Construct context-grounded prompt templates that inject external documentation into generation calls.
- Configure vector database schemas to enforce dimensional alignment between embedding models and vector indexes.
Detailed Concept Walkthrough
1. LLM Knowledge Boundaries and RAG Architecture
Standard LLMs possess frozen knowledge constrained by training cutoffs and lack access to proprietary runtime data, resulting in hallucinations when answering out-of-domain questions. RAG addresses this by dynamically querying external document stores at runtime and injecting relevant context directly into the prompt without model retraining.
- Mechanism: When a user submits a query, the RAG system first passes the question to a retrieval engine rather than directly to the generative LLM. The retrieval engine searches an external knowledge repository for relevant text snippets.
- Under the Hood: External data sources (PDFs, text files, APIs) are pre-processed and indexed into vector space. At query time, retrieved snippets are packed alongside system prompt boundaries into the context window of the LLM.
- Best Practice: Context boundary instructions must explicitly command the LLM to output a standardized lack-of-context notice if the retrieved text does not contain the answer, preventing fallback hallucinations.
# Grounding prompt template enforcing strict context boundaries
system_prompt = (
"You are a specialized assistant. Answer the user query using ONLY the provided context.\n"
"If the answer cannot be determined from the context, reply with: 'I lack sufficient context.'\n\n"
f"Context:\n{retrieved_context_chunk}\n\n"
f"User Query: {user_query}"
)
Key Takeaway: RAG replaces static knowledge recall with dynamic runtime context injection, eliminating the need for expensive model retraining.
2. Vector Embeddings and Dimensionality Matching
Embeddings convert unstructured text tokens into fixed-length arrays of continuous floating-point numbers that capture semantic meaning. Each model maps text into a specific vector dimension that must remain strictly consistent across all indexing and querying steps.
- Mechanism: Text passages are tokenized and processed through an embedding neural network, outputting a high-dimensional dense vector representing semantic concepts.
- Under the Hood: Models produce distinct fixed-size vector arrays; for instance, Ollama's
nomic-embed-textoutputs 768 dimensions, OpenAI'stext-embedding-3-smalloutputs 1536 dimensions, andtext-embedding-3-largeoutputs 3072 dimensions. - Syntax Rule: The database column or index dimension schema must match the exact output dimension of the active embedding model, or insertion and query operations will throw fatal dimension mismatch errors.
import chromadb
from chromadb.utils import embedding_functions
# Configure embedding function matching exact model dimensionality (e.g., 768 or 1536)
emb_fn = embedding_functions.DefaultEmbeddingFunction()
client = chromadb.Client()
collection = client.create_collection(name="japan_travel", embedding_function=emb_fn)
Key Takeaway: Embedding models output fixed vector dimensions that must identically match database column definitions and search query vectors.
3. Vector Databases and Semantic Similarity Retrieval
Vector databases store dense numerical vectors and execute high-speed nearest-neighbor calculations (such as cosine similarity) to find the most relevant context chunks for a given query vector.
- Mechanism: Raw documents are chunked, vectorized, and persisted into vector engines like Chroma DB, Pinecone, Weaviate, or PostgreSQL with pgvector. At search time, user queries are vectorized using the identical embedding model.
- Under the Hood: The vector engine computes distance metrics (e.g., cosine similarity or dot product) between the query vector and indexed document vectors, ranking chunks by semantic closeness rather than exact keyword overlap.
- Best Practice: Always isolate vector collections by domain and chunk size to prevent top-k retrieval dilution when querying dense documentation.
# Ingest documents and query vector database for top nearest neighbors
collection.add(
documents=["Tokyo transportation relies heavily on JR rail.", "Kyoto is famous for historic shrines."],
ids=["doc1", "doc2"]
)
results = collection.query(query_texts=["How do trains work in Tokyo?"], n_results=1)
top_chunk = results["documents"][0][0]
Key Takeaway: Semantic retrieval matches query intent to context using vector distance calculations rather than brittle literal keyword matches.
4. Local RAG Pipeline Execution with Ollama
Running local RAG pipelines with tools like Ollama enables air-gapped, privacy-compliant document retrieval and generation directly on local infrastructure without external API dependencies.
- Mechanism: Local models run inference locally via lightweight daemon processes, exposing HTTP endpoints for embedding generation and conversational completion.
- Execution Flow: A local script vectorizes a custom dataset using
nomic-embed-text, stores vectors in an in-memory or embedded store, and calls a local generator LLM with injected context to answer domain-specific questions. - Best Practice: Use automated lab provisioning to standardize library versions (Hugging Face transformers, Chroma DB, Flask) across local and enterprise deployment environments.
# Pull and serve local models using Ollama CLI
# CLI: ollama pull nomic-embed-text
# CLI: ollama run llama3
# Execute python retrieval pipeline
# CLI: python3 question.py --query "Explain container startup errors in Tokyo datacenter"
Key Takeaway: Local RAG pipelines provide fully private, reproducible context retrieval without incurring recurring API operational costs.
Topics Covered in Retrieval-Augmented Generation (RAG) Foundations
- LLM Knowledge Boundaries (0:00 - 1:50) — Examines static pretraining cutoffs, hallucinations, and why external retrieval is needed for specialized data.
- RAG Architecture Overview (1:50 - 2:34) — Traces the query-retrieval-generation lifecycle and runtime context injection workflow.
- Local Ollama Demonstration (2:34 - 4:35) — Runs a local Python RAG script against custom datasets to test in-domain and out-of-domain prompt boundaries.
- Embeddings and Dimensionality (4:35 - 7:32) — Explains text vectorization mechanics and vector dimension requirements across popular embedding models.
- Vector Databases and Search (7:32 - 9:22) — Details semantic similarity calculations, nearest-neighbor searches, and storage engines like Chroma DB and pgvector.
- Pipeline Deployment Setup (9:22 - 10:39) — Provisions the lab environment with Chroma DB, Hugging Face transformers, Flask, and OpenAI client libraries.
LLMs & Generative AI for Practitioners Cheat Sheet
-
ollama pull nomic-embed-text— Downloads local 768-dimension embedding model via Ollama CLIollama pull nomic-embed-text -
chromadb.Client()— Instantiates an in-memory or embedded vector database clientclient = chromadb.Client() -
collection.add()— Vectorizes and persists text documents with unique IDscollection.add(documents=["Text data"], ids=["id_1"]) -
collection.query()— Performs semantic similarity search for nearest-neighbor document chunkscollection.query(query_texts=["Search query"], n_results=2) -
CREATE TABLE ... vector(dim)— Defines pgvector column matching specific embedding model output sizeCREATE TABLE chunks (id serial, vec vector(768)); -
python3 question.py— Executes local Python RAG pipeline to test queriespython3 question.py
Comparison Table
| Embedding Model | Vector Dimension | Primary Deployment Context |
|---|---|---|
| nomic-embed-text | 768 | Local private pipelines via Ollama |
| text-embedding-3-small | 1536 | Cost-effective general cloud applications |
| text-embedding-3-large | 3072 | High-accuracy enterprise semantic search |
Common Pitfalls
- Mistake: Changing the embedding model without re-indexing existing stored vectors. Avoid: Always recreate and re-embed the vector database when switching embedding models.
- Mistake: Querying the vector database using a different model than the one used for ingestion. Avoid: Use the exact same embedding model for both indexing and runtime search queries.
- Mistake: Mismatching database vector column sizes with embedding model dimensions. Avoid: Define database schema dimensions explicitly to match the chosen model output array length.
- Mistake: Permitting the LLM to guess when retrieval returns zero matching context chunks. Avoid: Prompt the model to return an explicit lack-of-context message when context is missing.
FAQs
- Why use RAG instead of fine-tuning an LLM on proprietary data? RAG allows immediate updates without retraining costs and guarantees verifiable source citations with lower hallucination risks.
- What happens if query vectors have different dimensions than indexed document vectors? The vector database will raise a dimension mismatch exception and immediately fail the similarity calculation.
- Can I store vectors in traditional relational databases? Yes, relational databases like PostgreSQL support vector storage and cosine distance indexing through extensions like pgvector.
- How does semantic search differ from standard keyword search? Semantic search matches mathematical conceptual proximity in vector space rather than requiring exact lexical word matches.