Three high-impact prompting techniques cover the bulk of production work. Each solves a specific problem.
Pattern 1 — Few-Shot Prompting
Give the model 2-5 input/output examples before the actual request.
Classify the sentiment of customer reviews.
Examples:
Review: "Loved the shoes, fit perfectly."
Sentiment: positive
Review: "Quality is okay but shipping took forever."
Sentiment: mixed
Review: "Worst purchase ever, returning immediately."
Sentiment: negative
Now classify:
Review: "Decent product but customer service was unhelpful."
Sentiment:
Why it works: the model pattern-matches your examples and produces outputs in the same format.
When to use:
- Tasks with non-obvious output format (custom labels, specific structures).
- Stylistic consistency (always use specific phrasing).
- Reducing variance in outputs.
Choosing examples:
- Diverse — cover the cases you'll see.
- Representative — match the distribution of real requests.
- Include edge cases — especially the ones you've seen the model handle wrong.
Limits: too many examples (10+) burns tokens for diminishing returns. 3-5 examples is the sweet spot.
Pattern 2 — Chain-of-Thought (CoT)
Ask the model to reason step-by-step before producing the final answer.
Question: A store has 50 apples on Monday. They sell 60% on Tuesday and receive 30 more on Wednesday. How many do they have now?
Think through this step by step before giving the final answer.
vs. zero-shot:
Question: A store has 50 apples on Monday. They sell 60% on Tuesday and receive 30 more on Wednesday. How many do they have now?
Answer:
The CoT version usually gets it right; the zero-shot version often fails on multi-step math.
Why it works: the model literally generates intermediate steps. Each step becomes part of the next token's context. Errors compound less.
Modern frontier models (Claude 4, GPT-5) often do CoT implicitly — they include reasoning blocks before final answers. For mid-tier models, you may need to prompt it explicitly.
When to use:
- Multi-step reasoning.
- Math problems.
- Logic puzzles.
- Code debugging.
- Anything where the model needs to "think before speaking".
When NOT to use:
- Simple factual recall ("What's the capital of France?") — adds latency without benefit.
- Pattern-matching tasks where the right answer is immediate.
Trade-off: CoT increases output tokens (cost + latency). For high-volume simple tasks, skip it.
Pattern 3 — Structured Output
Constrain the model to produce JSON, XML, or other structured formats.
Output a JSON object with the following schema:
{
"person_name": "string",
"email": "string",
"phone": "string or null",
"city": "string or null",
"confidence": 0.0 to 1.0
}
Extract from this text:
"Hi, my name is Anuj Saini, you can reach me at anuj@example.com..."
Modern model APIs (Claude, GPT, Gemini) support structured output mode — pass a JSON schema, get guaranteed valid JSON back. Use it.
When to use:
- Pipeline outputs (downstream systems will parse them).
- Extraction tasks.
- Classification with multiple fields.
- Any task where the response feeds into code.
Practical tip: even when the API supports schema enforcement, include the schema in the prompt as well. Belt and suspenders.
Combining patterns
Production prompts often use all three:
SYSTEM: You are an extraction assistant. Read the customer message and extract key information.
EXAMPLES (few-shot):
Message: "Hi, I'm John, john@x.com, NYC, just bought widget 5 days ago, broken."
Output:
{
"name": "John",
"email": "john@x.com",
"city": "NYC",
"purchase_days_ago": 5,
"issue": "broken product",
"extraction_confidence": 0.95
}
(more examples...)
INSTRUCTION:
For the message below, think through what fields are present, then output JSON matching the schema.
Schema:
{
"name": "string",
"email": "string",
"city": "string or null",
"purchase_days_ago": "integer or null",
"issue": "string",
"extraction_confidence": 0.0-1.0
}
Message: "Hi I'm Priya from Bangalore. Bought shoes last week, want to return them."
The CoT element ("think through") + few-shot pattern + schema = consistent, accurate extraction.
ReAct, Tree-of-Thought, and others
Beyond the basics:
-
ReAct (Reasoning + Acting) — interleave reasoning and tool use. The model decides "I need data" → calls a tool → gets result → continues reasoning. Module 2 lesson on tool use covers this.
-
Tree-of-Thought — explore multiple reasoning branches in parallel, pick the best. Mostly research-grade; rare in production.
-
Self-consistency — generate multiple outputs with high temperature, take the majority vote. Improves accuracy for hard problems at 5-10x cost.
For most production work, master few-shot + CoT + structured output. The rest is optimization on the margin.
Common pattern mistakes
- Few-shot examples too similar. Cover variety; show edge cases.
- CoT for trivial tasks. Wastes tokens.
- No structured output for pipelines. Brittle parsing.
- Examples that contradict instructions. Model gets confused.
- Asking for CoT but cutting it off with low max_tokens. Model can't finish reasoning.
Takeaway
Few-shot for format consistency. CoT for reasoning. Structured output for pipelines. Combine when appropriate. These three patterns + good prompt anatomy = solid production prompts.