There are three families of LLMs to know about in 2026. Each has trade-offs that map to specific use cases.
Family 1 — Closed frontier models
Behind APIs from Anthropic, OpenAI, Google.
- Claude 4 (Opus, Sonnet, Haiku) — Anthropic. Strong reasoning, good for code, strong safety.
- GPT-5, GPT-4o, GPT-4o-mini — OpenAI. Strong general-purpose, excellent ecosystem.
- Gemini 2.5 Pro, Flash — Google. Long context (1M+ tokens), strong multimodal.
Pros:
- Best frontier performance on hard tasks.
- Hosted infra; just call an API.
- Continuously improving without effort from you.
Cons:
- Per-token costs add up at scale.
- Vendor lock-in risks.
- Data sent to third-party servers (privacy/compliance issues for some domains).
- No control over deprecation.
When to use: most starting points. The bulk of production LLM apps are on closed frontier APIs.
Family 2 — Open weight models
Downloadable, runnable on your own infra.
- Llama 3 / 4 (Meta) — strong general-purpose, multiple sizes (8B, 70B, 405B).
- Mistral / Mixtral — French; efficient mixture-of-experts; strong for the size.
- Qwen 2.5 (Alibaba) — strong multilingual, competitive with frontier on many tasks.
- DeepSeek — Chinese; very strong code and math; cost-efficient.
- Gemma 2 (Google) — small models (2B, 9B, 27B).
Pros:
- Self-hosted = data never leaves your infra.
- No per-token API fees (just compute cost).
- Can fine-tune freely.
- No deprecation risk.
Cons:
- Hardware required (or use a hosting service like Together, Fireworks, Replicate).
- Frontier closed models are still meaningfully better on hardest tasks.
- You manage scaling, monitoring, deployment yourself.
When to use:
- Privacy/compliance requirements (healthcare, finance, government).
- High-volume applications where API costs would be prohibitive.
- Fine-tuning is critical to the use case.
- Edge / on-device deployment.
Family 3 — Small specialized models
Either small open-weight models or task-specific fine-tunes.
- Gemma 2 2B — runs on a laptop, surprisingly capable.
- Phi 3 (Microsoft) — small, strong reasoning.
- TinyLlama, MobileLLM — for edge/mobile.
- Domain-fine-tunes — your own LoRA on top of an open base.
Pros:
- Run on CPU or modest GPU.
- Very low cost per request.
- Predictable latency.
Cons:
- Limited capability — need narrow task scope.
- May require fine-tuning to be useful.
When to use:
- High-volume narrow tasks (classification, structured extraction).
- Edge/mobile deployment.
- Cost-sensitive batch processing.
The selection grid
| Use case | Recommended starting point |
|---|---|
| Chatbot, general assistant | Claude Sonnet or GPT-4o (closed mid-tier) |
| Complex reasoning, coding | Claude Opus or GPT-5 (closed frontier) |
| Long-document Q&A | Gemini 2.5 Pro (1M+ context) or Claude with 200K |
| High-volume classification | Fine-tuned Llama 3 8B or Mistral 7B |
| Sensitive data (healthcare, finance) | Self-hosted Llama 3 70B or Qwen 2.5 |
| Code generation in product | Claude Sonnet, GPT-4o, or DeepSeek (specialist) |
| Multimodal (image + text) | GPT-4o, Claude Sonnet, or Gemini |
| Multilingual (especially Asian languages) | Qwen 2.5 or Gemini |
| Embedding generation | text-embedding-3-large, BGE-M3, or jina-embeddings-v3 |
Cost benchmarks (rough, 2026)
Approximate input/output cost per 1M tokens:
| Model | Input | Output |
|---|---|---|
| GPT-5 / Claude 4 Opus | $5-15 | $25-75 |
| Claude 4 Sonnet / GPT-4o | $3 | $15 |
| GPT-4o-mini / Claude 4 Haiku | $0.15-0.30 | $0.60-1.50 |
| Llama 3 70B (via Together) | $0.90 | $0.90 |
| Llama 3 8B (via Together) | $0.20 | $0.20 |
| Self-hosted (compute only) | varies | varies |
Mid-tier closed models hit a sweet spot — competitive with frontier on most tasks, ~5-10x cheaper.
How to compare models for YOUR task
Vendor benchmarks are mostly useless. Your task is different.
Process:
- Pick 3 candidate models (e.g., GPT-4o-mini, Claude Sonnet, Llama 70B via Together).
- Build a small eval set — 30-100 representative examples with expected outputs.
- Run all three. Compare outputs by hand or with LLM-as-judge.
- Track cost per request alongside quality.
- Pick the best quality/cost combination.
Module 3 covers evaluation in detail.
Common selection mistakes
- Picking based on benchmark scores. They don't reflect your task.
- Always using the most expensive model. Mid-tier is often equivalent at 5-10x lower cost.
- Switching models based on the hype cycle. Pick by data and let the data drive changes.
- Choosing open-source for ego. Self-hosting has real overhead; choose for genuine reasons.
- Not considering deprecation. Closed models occasionally retire; have an exit plan.
Takeaway
Three families: closed frontier (most starting points), open weight (privacy / scale), small specialized (edge / cost-critical). Pick by task, constraints, and economics — not by brand or hype. Always benchmark on your own data.