Chain-of-Thought Prompting: Step-by-Step Reasoning
Master chain-of-thought prompting to improve LLM accuracy on complex tasks through step-by-step reasoning, zero-shot CoT, and structured thinking patterns.
What you'll learn
- ✓Why chain-of-thought prompting improves accuracy
- ✓Zero-shot CoT with a single trigger phrase
- ✓Few-shot CoT with worked examples
- ✓Structured CoT for production-safe parsing
- ✓When CoT helps and when it hurts
Prerequisites
- •Basic understanding of LLM prompting
The Problem CoT Solves
Large language models are pattern matchers, not calculators. When you ask a model to jump directly from a complex question to an answer, it often gets it wrong because it has no intermediate workspace. Chain-of-thought (CoT) prompting fixes this by asking the model to show its work before giving the final answer.
Consider this example without CoT:
Q: A store sells shirts for $25 each. On Monday they sold 40 shirts.
On Tuesday they ran a 20% discount and sold 60 shirts. What was
the total revenue for both days?
A: $2,200
That answer is wrong. The correct answer is $2,200 only if you ignore the discount. The model jumped to a conclusion. Now with CoT:
Q: A store sells shirts for $25 each. On Monday they sold 40 shirts.
On Tuesday they ran a 20% discount and sold 60 shirts. What was
the total revenue for both days?
A: Let me work through this step by step.
- Monday: 40 shirts x $25 = $1,000
- Tuesday discount price: $25 x 0.80 = $20 per shirt
- Tuesday revenue: 60 shirts x $20 = $1,200
- Total: $1,000 + $1,200 = $2,200
The total revenue is $2,200.
Wait, this time the arithmetic is right because the model actually computed the discount. The intermediate steps create checkpoints where each computation builds on the last.
How CoT Works Internally
Direct: Question ──────────────────────> Answer
(one inference step, high error rate)
CoT: Question ──> Step 1 ──> Step 2 ──> Step 3 ──> Answer
| | |
v v v
Each token generated conditions the
next token, carrying state forward Each word the model generates becomes part of its context for the next word. By writing intermediate steps, the model gives itself a scratchpad. This is why CoT works better on larger models: they can maintain coherence across longer reasoning chains.
Zero-Shot CoT
The simplest form of CoT requires no examples. You add a single phrase to your prompt.
import openai
client = openai.OpenAI()
def solve_with_cot(problem: str) -> str:
response = client.chat.completions.create(
model="gpt-4o",
messages=[
{"role": "user", "content": f"{problem}\n\nLet's think step by step."}
],
temperature=0
)
return response.choices[0].message.content
The phrase “Let’s think step by step” was shown in the Kojima et al. (2022) paper to dramatically improve accuracy on arithmetic, logic, and commonsense reasoning tasks. Other effective trigger phrases include:
- “Let’s work through this carefully.”
- “Let’s break this down.”
- “Think about this step by step before answering.”
- “First, let’s identify what we know, then solve.”
They all do the same thing: instruct the model to generate intermediate reasoning tokens before the final answer.
Few-Shot CoT
Zero-shot CoT relies on the model figuring out what “step by step” means. Few-shot CoT removes that ambiguity by showing the model exactly what the reasoning should look like.
FEW_SHOT_COT_PROMPT = """Solve each problem by showing your reasoning, then give the final answer.
Q: A farmer has 3 fields. The first has 20 cows, the second has twice
as many, and the third has 10 fewer than the second. How many cows total?
A: Field 1: 20 cows
Field 2: 20 x 2 = 40 cows
Field 3: 40 - 10 = 30 cows
Total: 20 + 40 + 30 = 90 cows
ANSWER: 90
Q: A recipe calls for 2 cups of flour per dozen cookies. You want to
make 30 cookies. How many cups of flour do you need?
A: 1 dozen = 12 cookies, needs 2 cups
30 cookies = 30/12 = 2.5 dozen
Flour needed: 2.5 x 2 = 5 cups
ANSWER: 5
Q: {question}
A:"""
def solve_few_shot_cot(question: str) -> str:
prompt = FEW_SHOT_COT_PROMPT.format(question=question)
response = client.chat.completions.create(
model="gpt-4o",
messages=[{"role": "user", "content": prompt}],
temperature=0
)
return response.choices[0].message.content
The examples teach the model three things: the reasoning format (short lines, one computation per line), the level of detail (show the arithmetic), and the output structure (end with “ANSWER:”).
Structured CoT for Production
In production, you need the reasoning and the answer in separate, parseable fields. Mix CoT with structured output.
import json
STRUCTURED_COT_PROMPT = """Solve the following problem. Return your response as JSON with exactly two keys:
- "reasoning": your step-by-step working (string)
- "answer": the final numeric answer only (number)
Do not include any text outside the JSON object.
Problem: {problem}"""
def solve_structured(problem: str) -> dict:
response = client.chat.completions.create(
model="gpt-4o",
messages=[
{"role": "user", "content": STRUCTURED_COT_PROMPT.format(problem=problem)}
],
temperature=0,
response_format={"type": "json_object"}
)
result = json.loads(response.choices[0].message.content)
return result
# Example usage
result = solve_structured(
"A car travels 60 mph for 2.5 hours, then 45 mph for 1.5 hours. Total distance?"
)
# result = {
# "reasoning": "Phase 1: 60 mph x 2.5 hours = 150 miles. Phase 2: 45 mph x 1.5 hours = 67.5 miles. Total: 150 + 67.5 = 217.5 miles.",
# "answer": 217.5
# }
This pattern gives you the best of both worlds: the model reasons through the problem (improving accuracy), and your code gets a clean number to work with.
CoT for Non-Math Tasks
CoT is not just for arithmetic. It improves any task that involves multi-step reasoning.
Classification with explanation:
CLASSIFICATION_COT = """Classify the following customer message into one of these categories:
billing, technical, feature_request, complaint, praise
Before classifying, analyze the key phrases and sentiment in the message.
Message: "{message}"
Analysis:"""
# The model will first analyze key phrases, then classify.
# This catches nuanced messages that direct classification misses.
Code debugging:
DEBUG_COT = """The following Python function has a bug. Find it.
def find_duplicates(lst): seen = set() duplicates = [] for item in lst: if item in seen: duplicates.append(item) seen.add(item) return list(set(duplicates))
Think through what this function does step by step:
1. Trace through an example input
2. Identify where the behavior diverges from the expected output
3. Explain the bug
4. Show the fix"""
Decision making:
DECISION_COT = """Should we cache this API response?
Endpoint: /api/products/{id}
Average response time: 450ms
Requests per minute: 2,000
Data changes: every 4 hours
Response size: 12KB
Think through the tradeoffs:
1. What is the current cost (latency x volume)?
2. What would caching save?
3. What are the risks of stale data?
4. Recommendation with justification."""
When CoT Hurts
CoT is not always the right choice. It adds latency and cost because the model generates more tokens.
Skip CoT when:
- The task is simple lookup or classification with clear categories
- Latency is critical and the task does not require reasoning
- The model is small (under 7B parameters) and generates incoherent reasoning
- You are doing creative generation where “thinking” constrains creativity
Use CoT when:
- The task involves multi-step logic, math, or planning
- Accuracy matters more than speed
- You need to audit why the model gave a particular answer
- The problem has dependencies between parts
CoT Variants
Self-consistency: Run the same CoT prompt multiple times and take the majority answer. This improves accuracy on ambiguous problems at the cost of multiple API calls.
def solve_with_self_consistency(problem: str, n: int = 5) -> str:
answers = []
for _ in range(n):
response = client.chat.completions.create(
model="gpt-4o",
messages=[{"role": "user", "content": f"{problem}\n\nLet's think step by step."}],
temperature=0.7 # Need variation for self-consistency
)
text = response.choices[0].message.content
# Extract last number as the answer
numbers = [w for w in text.split() if w.replace('.','').replace(',','').isdigit()]
if numbers:
answers.append(numbers[-1])
# Return the most common answer
from collections import Counter
return Counter(answers).most_common(1)[0][0]
Plan-and-solve: Instead of “think step by step,” ask the model to first devise a plan, then execute it. This works well for complex multi-part problems.
Devise a plan to solve this problem, then execute each step of your plan.
Key Takeaways
Chain-of-thought prompting is one of the highest-leverage techniques in prompt engineering. Add “Let’s think step by step” for a quick zero-shot boost. Use few-shot examples when you need consistent reasoning format. Wrap reasoning in structured output for production systems. And know when to skip it: simple tasks do not benefit from forced reasoning, and small models generate unreliable chains.
Related articles
- Prompt Engineering Prompt Engineering: Self-Consistency for Reliable LLM Outputs
Learn how self-consistency prompting samples multiple reasoning paths and aggregates answers to improve accuracy, with hands-on examples and trade-offs.
- Prompt Engineering Prompt Evaluation: Measuring and Improving Quality
Learn how to measure prompt quality with evaluation datasets, scoring rubrics, A/B testing, and automated grading to iterate on prompts with evidence.
- Prompt Engineering Few-Shot Prompting: Learning from Examples
Master few-shot prompting to teach LLMs new tasks through carefully selected examples, formatting patterns, and example ordering strategies.
- Prompt Engineering Prompt Engineering for Code Generation
Learn prompt patterns for writing, reviewing, debugging, and refactoring code with LLMs, including practical templates and real examples.