Back to LLM Foundations: What They Actually Are

The Pretraining vs Fine-tuning vs Prompting Spectrum

Three ways to get an LLM to do what you want. Most teams pick the wrong one first. FIND_VIDEO: search 'LLM pretraining fine-tuning prompting comparison' — recommended channel: Andrej Karpathy / DeepLearning.AI. Aim for 10 min or under.

9 minutesVideo LessonPDF notes
🎯 Free Guest Mode: You are learning for free. Sign in to save your completion progress and quiz answers.

Ready to continue?

Mark this lesson as complete when you're ready to proceed.

Key moments

  1. Prompting for Tone and Persona — LangChain system prompt templates are configured to steer the tone and behavioral boundaries of base model responses.
  2. Grounding Responses with RAG — External bank databases are connected to inject dynamic customer records directly into the LLM context window.
  3. Domain Adaptation via Fine-Tuning — Historical conversation logs are utilized to retrain models so they proactively offer account optimizations in domain workflows.
  4. Base vs Instruct Models and PEFT — Foundation checkpoints are contrasted with instruct variants and configured with parameter-efficient LoRA adapters.
  5. Architectural Trade-offs and Hybrid Deployment — Cost, latency, and accuracy metrics are evaluated across prompting, retrieval, and fine-tuning for unified production architectures.
PDF notes

Frequently asked questions

Why is fine-tuning unreliable for memorizing factual data?

Fine-tuning updates parametric memory, which struggles with precise recall, produces hallucinations, and cannot be updated dynamically when facts change.

What is the difference between a Base model and an Instruct model?

Base models are next-token predictors that simply autocomplete text, whereas Instruct models have undergone fine-tuning to follow user instructions conversationally.

How does LoRA prevent catastrophic forgetting during domain adaptation?

LoRA leaves the original foundation model parameters completely frozen and only adjusts the weights of separate, low-rank adapter matrices added to the network.

When should I use RAG instead of prompt engineering?

Use RAG when the required information exceeds the context window, requires live database queries, or contains private, user-specific records.

How was this lesson?

Your feedback helps us refine explanations and catch bugs.