Part of AI Agent Handbook
Teach AI agents to think through problems step by step using chain-of-thought reasoning.
Step-by-Step Problem Solving (Chain-of-Thought)
You've learned how to write clear prompts and use strategies like roles and examples to guide your AI agent. But what happens when you ask your agent a question that requires real thinking: recalling facts and working through a problem step by step?
Try this experiment. Ask a language model: "If a train leaves Chicago at 2 PM traveling 60 mph, and another train leaves St. Louis (300 miles away) at 3 PM traveling 75 mph toward Chicago, when do they meet?"
You might get an answer. But is it right? The model might jump straight to a conclusion without showing its work. And when the answer is wrong, you have no idea where the reasoning broke down.
Now try adding one simple phrase: "Let's think this through step by step."
Suddenly, the model shows its reasoning. It breaks down the problem, considers each piece, and works toward the answer methodically. This simple technique, called chain-of-thought reasoning, transforms how AI agents handle complex problems.
Why Reasoning Matters
Language models are excellent at pattern matching and generating text. They can recall facts, write coherently, and follow instructions. But complex problems require more than pattern matching. They require reasoning: breaking down a problem, considering relationships, and building toward a solution.
Without explicit guidance to reason, models often take shortcuts. They might pattern-match to similar problems they've seen in training and output an answer that looks plausible but is actually wrong. This is especially common with:
- Math problems: Where each step depends on the previous one
- Logic puzzles: Where you need to track multiple constraints
- Multi-step tasks: Where you must plan a sequence of actions
- Analytical questions: Where you need to weigh evidence and draw conclusions
The solution isn't a more powerful model (though that can help). The solution is teaching the model to think through problems explicitly, showing its work as it goes.
What Is Chain-of-Thought Reasoning?
Chain-of-thought (CoT) reasoning is simple: instead of asking the model to jump straight to an answer, you prompt it to explain its thinking step by step. You're essentially asking it to "show its work," just like a math teacher would require.
When you use chain-of-thought prompting, the model generates intermediate reasoning steps before arriving at a final answer. These steps serve two purposes:
-
They improve accuracy: By working through the problem explicitly, the model is less likely to make logical errors or skip important considerations.
-
They provide transparency: You can see how the model arrived at its answer, which helps you trust the result and debug when something goes wrong.
Think of it like the difference between asking someone "What's ?" versus "What's ? Show me how you calculated it." The second request produces an answer and a process you can verify.
The Magic Phrase: "Let's Think Step by Step"
The simplest way to trigger chain-of-thought reasoning is to add a phrase like "Let's think through this step by step" or "Let's solve this step by step" to your prompt. This small addition signals to the model that you want explicit reasoning, not just a final answer.
Example (GPT-5)
Let's see this in action with a simple word problem:
import os
from openai import OpenAI
client = OpenAI(api_key=os.getenv("OPENAI_API_KEY"))
## Without chain-of-thought
prompt_simple = """A restaurant has 23 tables. Each table has 4 chairs.
If 12 chairs are broken and removed, how many chairs are left?"""
## With chain-of-thought
prompt_cot = """A restaurant has 23 tables. Each table has 4 chairs.
If 12 chairs are broken and removed, how many chairs are left?
Let's think through this step by step."""
## Get both responses
## Using GPT-5 for basic prompting and text generation
response_simple = client.chat.completions.create(
model="gpt-5", messages=[{"role": "user", "content": prompt_simple}]
)
response_cot = client.chat.completions.create(
model="gpt-5", messages=[{"role": "user", "content": prompt_cot}]
)
print("Without CoT:")
print(response_simple.choices[0].message.content)
print("\nWith CoT:")
print(response_cot.choices[0].message.content)Without CoT: Total chairs = 23 × 4 = 92. After removing 12 broken chairs: 92 − 12 = 80 chairs left. With CoT: - Total chairs initially: 23 tables × 4 chairs/table = 92 - Subtract broken chairs: 92 − 12 = 80 So, 80 chairs are left.
The first response might just say "80 chairs" (which is wrong, by the way). The second response will show the reasoning:
Let's think through this step by step.
Step 1: Calculate the total number of chairs
- 23 tables $\times$ 4 chairs per table = 92 chairs
Step 2: Subtract the broken chairs
- 92 chairs - 12 broken chairs = 80 chairs
Therefore, there are 80 chairs left in the restaurant.Wait, that's still 80. Let me recalculate: , then . Actually, that's correct! The point is that with chain-of-thought, you can verify each step. If there were an error, you'd see exactly where it occurred.
When Chain-of-Thought Helps Most
Chain-of-thought reasoning isn't necessary for every task. If you're asking for a fact ("What's the capital of France?") or a simple transformation ("Translate this to Spanish"), you don't need explicit reasoning steps. The model can answer directly.
But chain-of-thought becomes valuable when:
The problem has multiple steps: Each step builds on the previous one, and skipping a step leads to errors.
The answer isn't obvious: The problem requires analysis, not just recall.
You need to verify the logic: Being able to check the reasoning is as important as getting the right answer.
The model tends to make mistakes: If you notice the model getting certain types of questions wrong, adding chain-of-thought often helps.
Example: A Tricky Riddle
Riddles are a great test case because they require careful thinking. Let's try one:
Example (GPT-5)
from openai import OpenAI
client = OpenAI(api_key=os.getenv("OPENAI_API_KEY"))
riddle = """I speak without a mouth and hear without ears.
I have no body, but I come alive with wind. What am I?
Let's think through this step by step."""
## Using GPT-5 for reasoning through riddles and logic puzzles
response = client.chat.completions.create(
model="gpt-5", messages=[{"role": "user", "content": riddle}]
)
print(response.choices[0].message.content)An echo. It “speaks” by repeating sounds without a mouth, “hears” by receiving them without ears, has no body, and it exists only as sound traveling through air—so it metaphorically “comes alive” with the wind.
The model might respond:
Let's think through this step by step.
Step 1: Analyze the first clue - "I speak without a mouth"
This suggests something that makes sound but isn't alive or doesn't have a physical mouth.
Step 2: Consider the second clue - "hear without ears"
This means it can respond to or carry sound without having ears. This reinforces that it's not a living thing.
Step 3: Examine the third clue - "I have no body"
This confirms it's not a physical creature or person.
Step 4: Look at the final clue - "I come alive with wind"
This is the key clue. Something that needs wind to function and creates sound.
Step 5: Put it together
Something without a body that makes sound and requires wind... this describes an echo. Sound travels through air (wind), bounces off surfaces, and returns as an echo. An echo "speaks" (produces sound) without a mouth, "hears" (responds to sound) without ears, has no physical body, and exists because of sound waves traveling through air.
Answer: An echoBy working through each clue systematically, the model arrives at the correct answer with clear reasoning you can follow.
Teaching Through Examples: Few-Shot Chain-of-Thought
Remember few-shot prompting from the previous chapter? You can combine that technique with chain-of-thought by showing the model examples of step-by-step reasoning. This is especially powerful for tasks where you want consistent reasoning patterns.
Instead of showing input-output pairs, you show input-reasoning-output triplets. The model learns what to answer and how to think about the problem.
Example (GPT-5)
Let's say you're building a feature where your assistant helps evaluate whether claims are supported by evidence:
from openai import OpenAI
client = OpenAI(api_key=os.getenv("OPENAI_API_KEY"))
prompt = """Determine if the claim is supported by the evidence. Show your reasoning.
Example 1:
Claim: "Exercise improves mental health"
Evidence: "A study of 1,000 participants found that those who exercised 30 minutes daily reported 25% lower anxiety levels than those who didn't exercise."
Reasoning:
- The evidence comes from a study with a large sample size (1,000 participants)
- It shows a specific, measurable benefit (25% lower anxiety)
- Anxiety is a component of mental health
- The evidence directly relates to the claim
Conclusion: Supported
Example 2:
Claim: "Coffee causes heart disease"
Evidence: "Some people who drink coffee have reported heart palpitations."
Reasoning:
- The evidence is anecdotal ("some people reported")
- Heart palpitations are not the same as heart disease
- No causal relationship is established (correlation vs causation)
- The evidence is too weak to support the strong claim
Conclusion: Not supported
Now evaluate this:
Claim: "Reading before bed improves sleep quality"
Evidence: "A survey found that 60% of people who read before bed felt they slept better."
"""
## Using GPT-5 for analytical reasoning with few-shot examples
response = client.chat.completions.create(
model="gpt-5", messages=[{"role": "user", "content": prompt}]
)
print(response.choices[0].message.content)Reasoning: - The evidence is a self-reported survey; sample size and methodology are unspecified, raising concerns about reliability and bias. - It lacks a control or comparison group (e.g., people who don’t read before bed), so we can’t infer that reading caused the improvement. - “Felt they slept better” is subjective and may not reflect objective sleep quality; it also leaves 40% who did not report improvement. - The claim implies a general causal effect, while the evidence only shows that a subset of readers perceive a benefit, without establishing causation. Conclusion: Not supported. The evidence is insufficient to substantiate the causal claim.
The model will follow the reasoning pattern you demonstrated:
Reasoning:
- The evidence comes from a survey, which captures self-reported data
- 60% is a majority, suggesting a notable correlation
- "Felt they slept better" is subjective, not an objective measure of sleep quality
- The evidence shows correlation but doesn't prove causation (other factors could be involved)
- The sample size and methodology aren't specified, which limits confidence
Conclusion: Partially supported (shows correlation but not causation)By providing examples of good reasoning, you've taught the model how to approach this type of analysis.
Practical Applications for Your Personal Assistant
Let's apply chain-of-thought reasoning to make your personal assistant more capable. Here are some scenarios where it helps:
Planning a Multi-Step Task
from openai import OpenAI
client = OpenAI(api_key=os.getenv("OPENAI_API_KEY"))
prompt = """I need to prepare for a presentation next Tuesday. I need to:
- Research the topic (3 hours)
- Create slides (4 hours)
- Practice presenting (2 hours)
- Get feedback from a colleague (1 hour)
Today is Thursday. I have 2 hours available each evening (Thu, Fri, Mon).
I have 6 hours available on Saturday.
Create a schedule for completing these tasks. Think through this step by step."""
## Using GPT-5 for multi-step planning and scheduling
response = client.chat.completions.create(
model="gpt-5", messages=[{"role": "user", "content": prompt}]
)
print(response.choices[0].message.content)Here’s a clear plan that finishes everything with buffer and puts feedback late enough to be useful but early enough to act on it. Overall order (reasoning in brief): - Do research before heavy slide-building. - Finish a solid slide draft early Saturday, then practice so your colleague can react to both content and delivery. - Get feedback Monday and still have time to incorporate it. Schedule - Thu evening (2h) - Research: 2h - Goal: finalize outline, key message, and source list. - Fri evening (2h) - Research: 1h (completes 3h total) - Slides: 1h (build template, title/agenda, structure per outline) - Sat (6h total) - Slides: 3h (complete main content, charts, examples, speaker notes) - Practice: 2h (two focused run-throughs; time yourself; note rough spots) - Polish: 1h (tighten slides based on practice notes, refine transitions) - Mon evening (2h) - Feedback with colleague: 1h (share slides beforehand; do a brief run-through + Q&A) - Incorporate feedback: 1h (adjust slides, notes, and any tricky transitions) What this covers - Research: 3h (Thu 2h + Fri 1h) - Create slides: 4h (Fri 1h + Sat 3h) - Practice: 2h (Sat 2h) - Feedback: 1h (Mon 1h) - Buffer/improvements: 2h (Sat 1h polish + Mon 1h revisions) Quick prep actions - Today: book your colleague for Monday evening and send your outline by Friday night. - After Saturday practice: note top 3 issues to specifically ask your colleague about. - Optional: do a 10–15 min confidence run Tuesday morning if time allows. If your colleague can’t meet Monday evening, swap the Saturday “polish” hour with the feedback session on Saturday and use Monday’s 2h entirely for revisions and a final run-through.
The model will break down the problem:
Let's think through this step by step.
Step 1: Calculate total time needed
- Research: 3 hours
- Create slides: 4 hours
- Practice: 2 hours
- Feedback: 1 hour
Total: 10 hours
Step 2: Calculate available time
- Thursday evening: 2 hours
- Friday evening: 2 hours
- Saturday: 6 hours
- Monday evening: 2 hours
Total: 12 hours (enough time with 2 hours buffer)
Step 3: Consider dependencies
- Must research before creating slides
- Must create slides before practicing
- Should get feedback after practicing
Step 4: Create schedule
Thursday evening (2 hours): Research (part 1)
Friday evening (2 hours): Research (part 2, complete)
Saturday (6 hours): Create slides (4 hours) + Practice (2 hours)
Monday evening (2 hours): Get colleague feedback (1 hour) + final practice (1 hour)
This schedule respects the dependencies and uses 11 of your 12 available hours.Debugging Logic
Chain-of-thought is also valuable for helping users debug their own thinking:
from openai import OpenAI
client = OpenAI(api_key=os.getenv("OPENAI_API_KEY"))
prompt = """I'm trying to decide if I should buy a new laptop. Help me think through this decision step by step.
Current laptop: 4 years old, still works but slow
New laptop cost: \$1,200
My budget: \$2,000 saved for tech upgrades
Other needs: Might need a new phone in 6 months (\$800)
Should I buy the laptop now?"""
## Using GPT-5 for structured decision-making and analysis
response = client.chat.completions.create(
model="gpt-5", messages=[{"role": "user", "content": prompt}]
)
print(response.choices[0].message.content)Let’s walk through it step by step and land on a clear recommendation. 1) How painful is the slowness? - If your current laptop is costing you >10–15 minutes a day or feels unreliable, the productivity gain and reduced stress from a new one can easily justify the cost over the next few years. - If it’s mostly “annoying but workable,” consider a cheap tune‑up first (see Step 4). 2) Budget math (don’t forget tax and extras) - Laptop listed price: $1,200 - After tax (example 8%): ~$1,296 - Possible extras you might actually need: $0–$200 (warranty, dock, case, adapter) - Your tech fund: $2,000 - Phone in ~6 months: $800 (+ tax ≈ $864 at 8%) What this means: A $1,200 laptop likely leaves you with ~$704 after tax, which is short of the ~$864 you’ll want for the phone. You’ll need to either: - Save an additional ~$160+ over the next six months, or - Buy a slightly cheaper laptop (target ≤$1,050 before tax), or - Plan to sell/trade‑in your old laptop to cover the gap. 3) Timing and deals - Big discount windows typically come soon (back‑to‑school/late summer, and major sales events). Waiting 4–10 weeks could shave $100–$300 off. - New model cycles also create discounts on last year’s configs that are still excellent. 4) Lower‑cost ways to speed up your current laptop (often $50–$200) - If you have a hard drive, upgrade to an SSD. Biggest single speed boost. - Bump RAM to 16 GB if you run many browser tabs/apps. - Clean install the OS and remove startup bloat; ensure you have 20%+ free disk space. - Replace an old battery if it’s throttling performance. These can buy you 6–18 months for a fraction of the price, keeping phone money intact. 5) Resale/trade‑in can close the gap - A 4‑year‑old working laptop can often fetch $100–$300 depending on brand/spec/condition. That alone could cover tax on the new laptop or the phone shortfall. 6) Decision rule (pick the scenario that fits you best) - Buy now if: - The slowness is materially hurting your work or reliability is questionable, AND - You can either save an extra ~$200 before the phone purchase, sell the old laptop, or find a deal that keeps your all‑in laptop cost ≤$1,150 after tax. - Wait (or try a cheap upgrade) if: - The laptop is slow but stable, AND - You need to preserve the full ~$864 for the phone and don’t want to count on new savings, AND/OR - You can wait for an upcoming sale to bring the laptop closer to ~$1,000–$1,100 before tax. My take based on your numbers - If you can comfortably add at least ~$200 to your tech fund over the next six months (or are fine selling the old laptop), it’s reasonable to buy now—especially if the slowness is costing you time daily. - If adding savings is uncertain and the slowness is tolerable, hold off 4–8 weeks for a sale or do a $100–$200 upgrade now, then reassess. If you want, tell me: - Your main tasks (web, Office, coding, design, gaming) - Current laptop specs (CPU, RAM, storage type: HDD vs SSD) - Any brand/OS preferences I can help you decide whether an upgrade will fix it and/or suggest specific new models that fit an all‑in target and keep the phone budget safe.
The model will help structure the decision:
Let's think through this step by step.
Step 1: Assess the urgency
- Current laptop still works, just slow
- This suggests want, not urgent need
- You can likely wait if needed
Step 2: Check the budget
- Total budget: \$2,000
- Laptop cost: \$1,200
- Phone cost (in 6 months): \$800
- Total needed: \$2,000
- Remaining after both: \$0
Step 3: Consider timing
- If you buy the laptop now: \$800 left for phone
- That's exactly enough for the phone
- But no buffer for unexpected expenses
Step 4: Evaluate alternatives
- Wait on laptop, buy phone when needed: keeps more flexibility
- Buy laptop now: commits most of your budget
- Fix/upgrade current laptop: might be cheaper option
Step 5: Recommendation
Given that your current laptop works (just slow) and you have a more predictable phone need in 6 months, consider:
- Option A: Wait on the laptop purchase until after you assess the phone situation
- Option B: Look into cheaper laptop upgrades (RAM, SSD) to extend current laptop life
- Option C: If the laptop significantly impacts your work/productivity, buy it now but plan carefully for the phone purchase
The decision depends on how much the slow laptop affects your daily life versus the importance of maintaining budget flexibility.This kind of structured thinking helps users make better decisions by explicitly considering all factors.
Common Patterns in Chain-of-Thought Reasoning
As you use chain-of-thought prompting, you'll notice certain patterns that work well for different types of problems:
For math problems: Break into calculation steps, show each intermediate result
For logic puzzles: List constraints, test possibilities, eliminate what doesn't work
For planning tasks: Identify requirements, check resources, sequence actions
For analytical questions: State the question, gather relevant facts, weigh evidence, draw conclusions
For decision-making: Define options, list pros and cons for each, compare, recommend
You don't need to specify these patterns in your prompt. Just asking for step-by-step thinking often triggers the appropriate pattern. But if the model isn't structuring its reasoning the way you want, you can provide an example that demonstrates the pattern you prefer.
Limitations and When Not to Use Chain-of-Thought
Chain-of-thought reasoning is powerful, but it's not always the right tool:
It's slower: Generating reasoning steps takes more time than jumping to an answer. For simple questions, this overhead isn't worth it.
It uses more tokens: More generated text means higher API costs. Use chain-of-thought when accuracy matters more than speed or cost.
It can be verbose: Sometimes you just want a quick answer, not a detailed explanation. Match the technique to your needs.
It doesn't guarantee correctness: Chain-of-thought improves accuracy, but the model can still make errors in its reasoning. Always verify critical results.
The key is knowing when the benefits (better accuracy, transparency, debuggability) outweigh the costs (time, tokens, verbosity).
Building Intuition
Start by applying chain-of-thought to problems where the model makes mistakes. If a simple prompt produces wrong answers, try adding "Let's think step by step." You'll quickly notice which types of problems benefit most from explicit reasoning.
Keep track of what works. When you find a prompt pattern that produces good reasoning for a particular type of problem, save it. Over time, you'll build a library of effective chain-of-thought prompts you can reuse and adapt.
Pay attention to how the model structures its reasoning. You'll start recognizing good reasoning patterns versus sloppy ones. This helps you craft better prompts and evaluate the model's outputs more effectively.
Looking Ahead
Chain-of-thought reasoning is your first tool for teaching agents to think, not just respond. By prompting the model to show its work, you get more accurate answers and insight into how it arrived at them.
But we can go further. In the next chapter, you'll learn how to make your agent check its own work and refine its answers. You'll discover techniques for getting the agent to review its reasoning, consider alternatives, and improve its responses through self-reflection. These approaches build on chain-of-thought to create even more reliable agents.
The key takeaway: when you need your agent to handle complex problems, don't just ask for an answer. Ask it to think through the problem step by step. That simple change gives the model room to reason instead of jumping straight to a response.
Key Takeaways
- Chain-of-thought reasoning improves accuracy by making the model show its work instead of jumping to conclusions
- The phrase "Let's think step by step" is often all you need to trigger explicit reasoning
- Use chain-of-thought for complex problems where accuracy matters more than speed
- Combine with few-shot prompting to teach specific reasoning patterns
- Verify the reasoning, not just the answer, to catch errors and build trust
- Save effective patterns to build a library of proven chain-of-thought prompts
With chain-of-thought reasoning in your toolkit, your AI agent can handle problems that require real thinking. The next chapter builds on this foundation by teaching your agent to check and refine its own reasoning.
Quiz
Ready to test your understanding? Take this quick quiz to reinforce what you've learned about chain-of-thought reasoning.
Chain-of-Thought Reasoning Quiz
Reference
Citation details
Cite or share this article.
Continue with the full handbook
This chapter is part of AI Agent Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.
Explore AI Agent HandbookStay up to date
Get articles, book updates, and news delivered to your inbox.
No spam, unsubscribe anytime.
Join the community
Sign in to remove popups, track your reading progress, and join the discussion.

Comments
No comments yet. Be the first to share your thoughts!