Skip to content
Codeloom
Prompt Engineering

Chain-of-Thought Prompting: Step-by-Step Reasoning

Master chain-of-thought prompting to improve LLM accuracy on complex tasks through step-by-step reasoning, zero-shot CoT, and structured thinking patterns.

·8 min read · By Codeloom
Beginner 13 min read

What you'll learn

  • ✓Why chain-of-thought prompting improves accuracy
  • ✓Zero-shot CoT with a single trigger phrase
  • ✓Few-shot CoT with worked examples
  • ✓Structured CoT for production-safe parsing
  • ✓When CoT helps and when it hurts

Prerequisites

  • •Basic understanding of LLM prompting

The Problem CoT Solves

Large language models are pattern matchers, not calculators. When you ask a model to jump directly from a complex question to an answer, it often gets it wrong because it has no intermediate workspace. Chain-of-thought (CoT) prompting fixes this by asking the model to show its work before giving the final answer.

Consider this example without CoT:

Q: A store sells shirts for $25 each. On Monday they sold 40 shirts.
On Tuesday they ran a 20% discount and sold 60 shirts. What was
the total revenue for both days?

A: $2,200

That answer is wrong. The correct answer is $2,200 only if you ignore the discount. The model jumped to a conclusion. Now with CoT:

Q: A store sells shirts for $25 each. On Monday they sold 40 shirts.
On Tuesday they ran a 20% discount and sold 60 shirts. What was
the total revenue for both days?

A: Let me work through this step by step.
- Monday: 40 shirts x $25 = $1,000
- Tuesday discount price: $25 x 0.80 = $20 per shirt
- Tuesday revenue: 60 shirts x $20 = $1,200
- Total: $1,000 + $1,200 = $2,200

The total revenue is $2,200.

Wait, this time the arithmetic is right because the model actually computed the discount. The intermediate steps create checkpoints where each computation builds on the last.

How CoT Works Internally

Direct:  Question ──────────────────────> Answer
       (one inference step, high error rate)

CoT:     Question ──> Step 1 ──> Step 2 ──> Step 3 ──> Answer
                    |          |          |
                    v          v          v
                 Each token generated conditions the
                 next token, carrying state forward
Direct prompting vs chain-of-thought

Each word the model generates becomes part of its context for the next word. By writing intermediate steps, the model gives itself a scratchpad. This is why CoT works better on larger models: they can maintain coherence across longer reasoning chains.

Zero-Shot CoT

The simplest form of CoT requires no examples. You add a single phrase to your prompt.

import openai

client = openai.OpenAI()

def solve_with_cot(problem: str) -> str:
    response = client.chat.completions.create(
        model="gpt-4o",
        messages=[
            {"role": "user", "content": f"{problem}\n\nLet's think step by step."}
        ],
        temperature=0
    )
    return response.choices[0].message.content

The phrase “Let’s think step by step” was shown in the Kojima et al. (2022) paper to dramatically improve accuracy on arithmetic, logic, and commonsense reasoning tasks. Other effective trigger phrases include:

  • “Let’s work through this carefully.”
  • “Let’s break this down.”
  • “Think about this step by step before answering.”
  • “First, let’s identify what we know, then solve.”

They all do the same thing: instruct the model to generate intermediate reasoning tokens before the final answer.

Few-Shot CoT

Zero-shot CoT relies on the model figuring out what “step by step” means. Few-shot CoT removes that ambiguity by showing the model exactly what the reasoning should look like.

FEW_SHOT_COT_PROMPT = """Solve each problem by showing your reasoning, then give the final answer.

Q: A farmer has 3 fields. The first has 20 cows, the second has twice
as many, and the third has 10 fewer than the second. How many cows total?
A: Field 1: 20 cows
Field 2: 20 x 2 = 40 cows
Field 3: 40 - 10 = 30 cows
Total: 20 + 40 + 30 = 90 cows
ANSWER: 90

Q: A recipe calls for 2 cups of flour per dozen cookies. You want to
make 30 cookies. How many cups of flour do you need?
A: 1 dozen = 12 cookies, needs 2 cups
30 cookies = 30/12 = 2.5 dozen
Flour needed: 2.5 x 2 = 5 cups
ANSWER: 5

Q: {question}
A:"""

def solve_few_shot_cot(question: str) -> str:
    prompt = FEW_SHOT_COT_PROMPT.format(question=question)
    response = client.chat.completions.create(
        model="gpt-4o",
        messages=[{"role": "user", "content": prompt}],
        temperature=0
    )
    return response.choices[0].message.content

The examples teach the model three things: the reasoning format (short lines, one computation per line), the level of detail (show the arithmetic), and the output structure (end with “ANSWER:”).

Structured CoT for Production

In production, you need the reasoning and the answer in separate, parseable fields. Mix CoT with structured output.

import json

STRUCTURED_COT_PROMPT = """Solve the following problem. Return your response as JSON with exactly two keys:
- "reasoning": your step-by-step working (string)
- "answer": the final numeric answer only (number)

Do not include any text outside the JSON object.

Problem: {problem}"""

def solve_structured(problem: str) -> dict:
    response = client.chat.completions.create(
        model="gpt-4o",
        messages=[
            {"role": "user", "content": STRUCTURED_COT_PROMPT.format(problem=problem)}
        ],
        temperature=0,
        response_format={"type": "json_object"}
    )
    result = json.loads(response.choices[0].message.content)
    return result

# Example usage
result = solve_structured(
    "A car travels 60 mph for 2.5 hours, then 45 mph for 1.5 hours. Total distance?"
)
# result = {
#     "reasoning": "Phase 1: 60 mph x 2.5 hours = 150 miles. Phase 2: 45 mph x 1.5 hours = 67.5 miles. Total: 150 + 67.5 = 217.5 miles.",
#     "answer": 217.5
# }

This pattern gives you the best of both worlds: the model reasons through the problem (improving accuracy), and your code gets a clean number to work with.

CoT for Non-Math Tasks

CoT is not just for arithmetic. It improves any task that involves multi-step reasoning.

Classification with explanation:

CLASSIFICATION_COT = """Classify the following customer message into one of these categories:
billing, technical, feature_request, complaint, praise

Before classifying, analyze the key phrases and sentiment in the message.

Message: "{message}"

Analysis:"""

# The model will first analyze key phrases, then classify.
# This catches nuanced messages that direct classification misses.

Code debugging:

DEBUG_COT = """The following Python function has a bug. Find it.

def find_duplicates(lst): seen = set() duplicates = [] for item in lst: if item in seen: duplicates.append(item) seen.add(item) return list(set(duplicates))


Think through what this function does step by step:
1. Trace through an example input
2. Identify where the behavior diverges from the expected output
3. Explain the bug
4. Show the fix"""

Decision making:

DECISION_COT = """Should we cache this API response?

Endpoint: /api/products/{id}
Average response time: 450ms
Requests per minute: 2,000
Data changes: every 4 hours
Response size: 12KB

Think through the tradeoffs:
1. What is the current cost (latency x volume)?
2. What would caching save?
3. What are the risks of stale data?
4. Recommendation with justification."""

When CoT Hurts

CoT is not always the right choice. It adds latency and cost because the model generates more tokens.

Skip CoT when:

  • The task is simple lookup or classification with clear categories
  • Latency is critical and the task does not require reasoning
  • The model is small (under 7B parameters) and generates incoherent reasoning
  • You are doing creative generation where “thinking” constrains creativity

Use CoT when:

  • The task involves multi-step logic, math, or planning
  • Accuracy matters more than speed
  • You need to audit why the model gave a particular answer
  • The problem has dependencies between parts

CoT Variants

Self-consistency: Run the same CoT prompt multiple times and take the majority answer. This improves accuracy on ambiguous problems at the cost of multiple API calls.

def solve_with_self_consistency(problem: str, n: int = 5) -> str:
    answers = []
    for _ in range(n):
        response = client.chat.completions.create(
            model="gpt-4o",
            messages=[{"role": "user", "content": f"{problem}\n\nLet's think step by step."}],
            temperature=0.7  # Need variation for self-consistency
        )
        text = response.choices[0].message.content
        # Extract last number as the answer
        numbers = [w for w in text.split() if w.replace('.','').replace(',','').isdigit()]
        if numbers:
            answers.append(numbers[-1])
    # Return the most common answer
    from collections import Counter
    return Counter(answers).most_common(1)[0][0]

Plan-and-solve: Instead of “think step by step,” ask the model to first devise a plan, then execute it. This works well for complex multi-part problems.

Devise a plan to solve this problem, then execute each step of your plan.

Key Takeaways

Chain-of-thought prompting is one of the highest-leverage techniques in prompt engineering. Add “Let’s think step by step” for a quick zero-shot boost. Use few-shot examples when you need consistent reasoning format. Wrap reasoning in structured output for production systems. And know when to skip it: simple tasks do not benefit from forced reasoning, and small models generate unreliable chains.