Back to M2 — RAG Architecture + Evals

RAG Architecture + Trade-offs

Outcome: Choose RAG vs fine-tune for 3 cases Curated video (IBM Technology): What is Retrieval-Augmented Generation (RAG)? — https://www.youtube.com/watch?v=T-D1OfcDW1M (verified live via yt-dlp 2026-09-24). Pointer: llms-genai-for-practitioners/12 (hybrid legal-AI pattern); shell: courses/video-scripts/genai-rag-agents/03.md.

7 minutesVideo LessonPDF notes
🎯 Free Guest Mode: You are learning for free. Sign in to save your completion progress and quiz answers.

Ready to continue?

Mark this lesson as complete when you're ready to proceed.

Key moments

  1. RAG vs FT Overview — The lesson introduces the core decision point between updating knowledge via context (RAG) or updating model weights (FT).
  2. Data Freshness Constraint — Discusses how the speed of knowledge change dictates the preference for RAG due to its low update cost compared to retraining.
  3. Hybrid Retrieval Need — Explains why single-method retrieval fails and introduces the concept of combining sparse and dense search for better coverage.
  4. Fusion and Re-ranking — Details the mechanism of Reciprocal Rank Fusion (RRF) for merging and optimizing the results from the hybrid search pipeline.
  5. Grounded Generation — Focuses on the final step of the RAG pipeline, ensuring the LLM output is strictly attributable to the source documents to prevent hallucination.
  6. Case Study Decisions — Reviews three specific enterprise cases and applies the RAG/FT decision matrix based on data volatility and required output style.
PDF notes

Frequently asked questions

Does RAG eliminate the need for a large LLM?

No. RAG provides knowledge, but a larger LLM is still required for complex reasoning, synthesis, and instruction following capabilities.

When is fine-tuning absolutely necessary over RAG?

When you need the model to adopt a specific, non-standard output format, or when the task requires deep, internalized domain reasoning that cannot be easily provided as context.

What is the main trade-off of using hybrid retrieval?

Increased latency during the retrieval phase, as two separate search operations (sparse and dense) must be executed and their results fused.

How do I handle the 'Information not found' scenario in RAG?

Explicitly instruct the LLM in the system prompt to state that the answer is unavailable if the context does not contain the necessary information.

How was this lesson?

Your feedback helps us refine explanations and catch bugs.