You don't need to derive backpropagation to build with LLMs. You do need a mental model of what's happening underneath. Three concepts cover 80% of it.
Concept 1 — Tokens
LLMs don't process words; they process tokens. A token is a chunk of text — sometimes a whole word, sometimes part of a word, sometimes punctuation.
- "Hello" → 1 token.
- "indistinguishable" → 4-5 tokens (depending on tokenizer).
- "The cat sat" → 3-4 tokens.
- Code, JSON, non-English text → often more tokens per character.
Why this matters:
- Pricing is per-token. 1M tokens of input/output costs $X.
- Context windows are measured in tokens. Claude 4 has a 200K-token (or 1M-token for premium) context window. GPT-4o has 128K.
- Latency scales with token count.
Rough rule: 1 English word ≈ 1.3 tokens. 100 tokens ≈ 75 words.
Test it: OpenAI has a tokenizer at platform.openai.com/tokenizer; Anthropic at console.anthropic.com. Paste text, see token counts.
Concept 2 — Embeddings
The model converts each token into a vector — a list of ~1000-4000 numbers that encode the token's meaning in context.
Tokens with similar meaning have similar embeddings:
"king"and"queen"are close in embedding space."dog"and"puppy"are close."king"and"refrigerator"are far apart.
This is what lets the model "understand" relationships. Math on embeddings approximates math on meaning.
Embeddings power:
- RAG (retrieval-augmented generation): convert documents to embeddings, find similar ones at query time.
- Semantic search: search by meaning, not keywords.
- Clustering: group similar texts.
You'll use embeddings constantly even without thinking about the transformer itself.
Concept 3 — The Transformer (the simplified mental model)
The transformer is the architecture inside every modern LLM. Skip the math; here's what it does:
- Input: a sequence of tokens.
- For each token, predict the next one, given everything before it.
- Repeat to generate text.
The "attention" mechanism lets the model look back at any token in the context, not just the most recent. That's why LLMs handle long-range dependencies (referring back to something 50 tokens ago) gracefully.
Two practical implications:
Autoregressive generation
LLMs generate one token at a time, left-to-right. Each new token depends on all previous tokens. This explains:
- Latency: longer outputs take longer (linearly).
- Streaming: you can show tokens as they're generated (good UX).
- Why prompts work: the model literally bases the next token on what's already there.
Context window limits
The model can attend to all tokens in its context, but only up to the maximum window size. Past that, older content gets dropped or the call fails.
If you put 250K tokens into a 200K-window model, you're either truncated, rejected, or moved to a different tier.
What LLMs ARE good at
- Pattern completion (the obvious one — that's what they're trained for).
- Summarization, rewriting, translation.
- Following instructions in natural language.
- Reasoning over presented context (when context fits).
- Code generation, especially for well-known languages and patterns.
- Creative tasks: writing, brainstorming, ideation.
What LLMs ARE NOT good at
- Math beyond simple arithmetic (use tool calling — Module 2).
- Precise counting (asking "how many words in this article" — wrong half the time).
- Long-range factual recall without context (use RAG — Module 3).
- Following extremely complex multi-step instructions without prompting help.
- Anything they weren't trained on — recent events, your private data, unusual domains.
- Determinism. Same input → similar but not identical output (unless temperature=0).
The cost-quality-speed triangle
In 2026's model landscape:
- Frontier models (Claude 4 Opus, GPT-5, Gemini 2.5 Pro): best quality, highest cost, slowest.
- Mid-tier (Claude Sonnet, GPT-4o-mini, Llama 3 70B): great cost/performance ratio.
- Small/local (Llama 3 8B, Mistral 7B, Gemma): cheap, fast, good for narrow tasks.
You'll likely use a mix — frontier for hard reasoning, mid-tier for the bulk of work, small for high-volume simple tasks.
What this lets you do
By the end of this lesson you should be able to:
- Estimate the token count of a prompt (and therefore its cost).
- Decide whether a task fits in a given context window.
- Pick a starting model tier based on task complexity.
- Recognize when a problem is suited to LLMs vs needs a different tool.
That's enough to build real LLM applications. The rest of this course is "what to do with that knowledge".
Common confusion
-
"LLMs know everything they were trained on." They've seen a lot, but they don't reliably remember specifics. Treat any factual claim as suspect; use RAG if accuracy matters.
-
"Higher temperature = smarter." No. Temperature controls randomness. 0 = deterministic and consistent; 1 = more varied and creative. Hard tasks usually want low temperature.
-
"GPT is better than Claude is better than..." Models leapfrog each other constantly. Build your code to be model-agnostic; benchmark on your actual task.
-
"Bigger context window = better." Not necessarily. Models often have a "lost in the middle" problem — they pay less attention to content in the middle of a long context. More context isn't always better.
Takeaway
LLMs predict the next token using transformer attention over embedded tokens. They're brilliant at pattern completion, terrible at precise calculation, and need help (RAG, tools) for facts and math. Pick the model tier by task complexity, not brand loyalty.