Back to LLM Foundations: What They Actually Are

Tokens, Context Windows, and Why They Matter

The fundamental constraint that shapes every LLM application: tokens in, tokens out, with a hard limit. FIND_VIDEO: search 'LLM context window tokens lost in middle' — recommended channel: AI Engineer / Greg Kamradt. Aim for 10 min or under.

10 minutesVideo LessonPDF notes
🎯 Free Guest Mode: You are learning for free. Sign in to save your completion progress and quiz answers.

Ready to continue?

Mark this lesson as complete when you're ready to proceed.

Key moments

  1. Context Anatomy and Limits — Explains how input and output tokens combine against provider limits and triggers hard execution errors.
  2. Benchmarking Context Sizes — Compares nominal context capacities across model architectures using tools like models.dev.
  3. Attention Mechanics and Retrieval Loss — Details how primacy and recency bias cause the Lost in the Middle phenomenon.
  4. Context Management in Claude Code — Demonstrates how to monitor tokens and manage state using context, clear, and compact.
  5. Mitigating Tool and Config Bloat — Explains the static context overhead introduced by MCP schemas and oversized rule files.
  6. Nominal Size vs Retrieval Reliability — Contrasts vendor context window claims with practical needle-in-a-haystack retrieval performance.
PDF notes

Frequently asked questions

Why does model performance degrade even if I stay below the hard token limit?

Due to attention mechanisms exhibiting primacy and recency bias, models struggle to recall details placed in the conversational middle. Bloated context disperses attention across non-essential tokens.

What is the difference between running /compact and running /clear?

The /clear command wipes history back to baseline at zero cost. The /compact command calls an LLM to summarize previous turns, preserving key context while consuming tokens and time.

How do MCP servers consume context before I send a prompt?

MCP servers inject complete tool definitions and parameter schemas directly into the prompt payload. This static overhead is resent on every turn regardless of whether tools are invoked.

When should I proactively compact or clear my context window?

Trigger a reset or compaction when remaining free tokens dip below safety margins like 50k tokens or when switching to an unrelated coding task.

How was this lesson?

Your feedback helps us refine explanations and catch bugs.