Last updated: 2026-09-28
How LLMs Reason, Self-Correct and Get Checked
A language model writes its answer one token at a time and can't take one back. That single fact explains most of what people call its "thinking," why it sometimes corrects itself mid-answer and often doesn't, and why the software wrapped around a model matters as much as the model does. This page follows the thread from the token to the agent: what a model does on its own, what extra "thinking" buys, where self-correction works and where it fails, and how the surrounding framework supplies feedback the model can't generate for itself, including checking the sources it cites.
One Token at a Time FoundationalKnowledge that endures for decades — core principles
At each step the model produces a probability for every token in its vocabulary, the system picks one, appends it to the text so far, and repeats. Understanding Large Language Models covers the mechanism; two details matter here. The pick isn't always the top-ranked token, because always choosing the most probable continuation tends to produce bland, repetitive text. Systems usually sample from the distribution instead, often trimming its unlikely tail first [1], which is why the same prompt can give different answers on different runs.
Two consequences follow, and the rest of the page builds on both. The text is append-only: the model has no backspace, and every token already written becomes context that shapes the next one. And nothing in the loop consults the world. The model's only input is the tokens in front of it.
What "Thinking" Adds Ephemeral / ToolingKnowledge that evolves in months to a year — check for updates
Asking a model to write out intermediate steps improves its performance on multi-step problems [2], a technique covered on the orchestration page. The steps do work. Each one is a set of tokens the model can build on, so a problem that needs an impossible single leap becomes a series of short ones, and every step spends more computation on the problem.steps are just more tokens
Reasoning models train that behaviour in. DeepSeek-R1, described in Nature in 2025, was trained with reinforcement learning rather than human-written reasoning examples, on maths, coding and science problems whose answers can be graded automatically. Its long reasoning traces came to include self-reflection, verification and switching strategy [3]. Snell and colleagues studied the same lever from the other side: in a compute-matched comparison, on problems where a smaller model already had some success, letting it spend more computation at answer time could beat a model fourteen times larger, and the best way to spend that computation depended on how hard the question was [4].
Simply put, "thinking" is more tokens of working before the answer, with more computation behind them. It isn't a separate faculty added alongside next-token prediction. How well a given model does it depends on how the model was trained, and the next section shows why that matters.
Going Back: Two Places Correction Can Live Applied / MethodologicalKnowledge with a 5–10 year half-life — stable practice
Because text is append-only, going back can't mean deleting. Inside a single response it means writing new text that overrides the old: a line such as "wait, 17 × 24 is 408, not 358," followed by a redo. The mistaken step stays in the context, and the model has to weigh the correction against it.
17 × 24 = 358"] --> B["Step 2:
358 + 40 = 398"] B --> C["Wait: 17 × 24
is 408, not 358"] C --> D["Redo:
408 + 40 = 448"] D --> E["Answer: 448"] style A fill:#F2B8B5 style B fill:#F2B8B5 style E fill:#8FBF6A
That kind of in-text correction is what the R1 traces contain [3], and it is a trained capability. A model without reasoning training can be told to double-check its work and will produce text that sounds like checking, without having been optimised to catch its own slips.
The other place correction can live is outside the model, in code that runs the model several times or steers it. Self-consistency samples several independent reasoning paths and takes the answer most of them reach, which lifted accuracy by up to 17.9 points on the GSM8K arithmetic benchmark [5]. Tree of Thoughts goes further: code around the model proposes several next steps, has the model rate them, drops the weak branches, and looks ahead or backtracks when a branch dead-ends. On the Game of 24 puzzle, GPT-4 with chain-of-thought prompting solved 4% of tasks and with Tree of Thoughts 74% [6]. The model never went back in its own text there. The search procedure around it did. Step-level feedback helps at training time too: Lightman and colleagues found that grading each step of a solution beat grading only the final answer for training models on MATH problems [7].
The distinction between correction in the model's own text and correction imposed by a procedure around it is our framing, not a term from those papers. It matters because the two fail differently.
When Self-Correction Fails Applied / MethodologicalKnowledge with a 5–10 year half-life — stable practice
A fluent correction isn't a reliable one. Huang and colleagues found that, on reasoning tasks, models asked to review and revise their own answers without any outside signal often didn't improve and sometimes got worse [8]. Kamoi and colleagues' survey of the field points the same way: self-correction works well where reliable external feedback exists, and large-scale fine-tuning helps, but they found no demonstration of success from prompting a model to critique itself, outside a few narrow settings [9].fluent isn't the same as checked
The reasoning on display has a second weakness. Turpin and colleagues changed something irrelevant to correctness, such as the order of multiple-choice options, and found that models' answers shifted while their written explanations didn't mention it [10]. Chen and colleagues tested reasoning models with hints planted in the prompt: the visible chain of thought acknowledged a hint the model had used in at least 1% of cases, and the rate was often below 20% [11]. Explaining What Language Models Do develops that faithfulness problem in full.
A reasoning trace is therefore evidence about how an answer might have arisen, and it isn't a log of how it did arise. A model's confident "I've double-checked this" shows nothing unless the check touched something outside the model.
What the Framework Around the Model Adds Applied / MethodologicalKnowledge with a 5–10 year half-life — stable practice
An agentic system is the model plus a harness: code that runs the loop, offers tools, feeds results back, and decides when to stop. Anthropic's guidance on building agents separates workflows, where code fixes the path, from agents, where the model directs its own steps and tool use, and says an agent needs "ground truth from the environment" at each step, such as tool results or code execution, to judge its progress. It also recommends stopping conditions such as a maximum number of iterations, and human checkpoints [12]. The loop at the core is the thought, action, observation cycle described on the orchestration page [13].
Each element the harness adds supplies something the model can't produce alone:
- An observation from the world. A test result, a search hit, the contents of a file.
- A limit. An iteration cap, a budget, a list of permitted tools.
- A second opinion. A separate evaluator, a rubric, or a person.
- Memory across attempts. Notes on what failed last time.
or a tool call"] --> T["Tool runs:
fetch, search, test, lookup"] T --> R["Result comes back
as ground truth"] R --> C{"Checks pass?"} C -->|yes| H["Human checkpoint
or final output"] C -->|no| L{"Retries left?"} L -->|yes| F["Feed the failure
back to the model"] F --> G H -.->|"optional: reviewer
sends it back"| F L -->|no| E["Stop and hand
to a person"] style H fill:#8FBF6A style E fill:#FFC857
Reflexion shows the pattern working. An agent attempts a task, receives a signal from the environment such as unit-test results, writes a short reflection on what went wrong into a memory buffer, and tries again with that reflection in context, with no change to the model's weights. The authors reported 91% pass@1 on the HumanEval coding benchmark, against 80% for GPT-4 at the time [14]. The improvement rests on a signal from the environment, which fits the survey's finding above.
It helps to picture three layers of feedback. This layering is our own framing, added on top of the papers cited. At the token layer, the sampling distribution decides each word. At the reasoning layer, the model's own trace can question earlier steps, unreliably. At the framework layer, code and tools compare the output with something outside the model. The further a check sits from the model's own text, the harder it is for the model to talk its way past it.
Tools That Fetch and Validate Sources Applied / MethodologicalKnowledge with a 5–10 year half-life — stable practice
Citations show the gap clearly. Walters and Wilder asked two ChatGPT versions for short literature reviews on 42 topics and checked all 636 citations. Of the GPT-3.5 citations, 55% were fabricated, against 18% of the GPT-4 citations. Of the citations that did exist, 43% (GPT-3.5) and 24% (GPT-4) contained substantive errors [15]. Those are 2023 models, and current ones will score differently. The mechanism still applies: a citation is a plausible string of tokens, and next-token prediction can produce a plausible string for a source that doesn't exist.plausible string, missing source
A tool loop can check the parts of a citation that are checkable:
| Check | Can code do it? | How |
|---|---|---|
| The URL or DOI exists | Yes, deterministically | Request the URL; resolve the DOI |
| Title, authors, venue and year match the claim | Yes | Compare against a bibliographic record such as Crossref's |
| The quoted passage or figure appears in the source | Yes | String match against the fetched text |
| The source supports the claim made | Not reliably | A model reading both can give a first pass; a person decides |
| The source is any good | No | Judgement about provenance and quality |
An agent with a fetch tool and a bibliographic lookup can run the first three checks on every reference and hand back only the ones that fail. That turns the model from an author who might invent a reference into a drafter whose references have been audited by something that fluent text can't charm. The fourth check still needs judgement. CRITIC showed that letting a model verify its output through tools such as search and code execution, then revise, improved results on question answering, maths and toxicity reduction [16]. The caution from the retrieval-augmented generation page still applies, though: a passage in context doesn't guarantee that the answer follows from it. For anything that will be published, a person makes the last call.
Tools bring a risk of their own. A fetched page is text the model reads, and text can carry instructions. Greshake and colleagues demonstrated indirect prompt injection, where instructions planted in content an LLM-integrated application retrieves steer what the application does [17]. Treat fetched content as data, limit which tools an agent may call, and keep the checks that matter in code that page text can't reach.
Related Topics
- Understanding Large Language Models — the next-token mechanism and the structural source of hallucination that this page builds on.
- LLM Orchestration, Context Engineering & Agentic AI — chain-of-thought prompting and the thought, action, observation loop in more detail.
- Retrieval-Augmented Generation — grounding an answer in retrieved passages, and why a passage in context still needs checking.
- Designing Auditable, Robust Agentic Systems — why verification has to come from outside the component being verified, and where to place human checkpoints.
- Explaining What Language Models Do: Methods and Their Limits — why a written chain of thought can differ from what drove the answer.
References
- Holtzman, A., Buys, J., Du, L., Forbes, M., & Choi, Y. (2020). The Curious Case of Neural Text Degeneration. International Conference on Learning Representations (ICLR 2020). arXiv:1904.09751
- Wei, J., Wang, X., Schuurmans, D., Bosma, M., Ichter, B., Xia, F., Chi, E., Le, Q., & Zhou, D. (2022). Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. Advances in Neural Information Processing Systems, 35, 24824–24837.
- Guo, D., Yang, D., Zhang, H., et al. (2025). DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning. Nature, 645, 633–638. https://doi.org/10.1038/s41586-025-09422-z
- Snell, C., Lee, J., Xu, K., & Kumar, A. (2025). Scaling LLM Test-Time Compute Optimally Can Be More Effective than Scaling Parameters for Reasoning. International Conference on Learning Representations (ICLR 2025). arXiv:2408.03314
- Wang, X., Wei, J., Schuurmans, D., Le, Q., Chi, E., Narang, S., Chowdhery, A., & Zhou, D. (2023). Self-Consistency Improves Chain of Thought Reasoning in Language Models. International Conference on Learning Representations (ICLR 2023). arXiv:2203.11171
- Yao, S., Yu, D., Zhao, J., Shafran, I., Griffiths, T. L., Cao, Y., & Narasimhan, K. (2023). Tree of Thoughts: Deliberate Problem Solving with Large Language Models. Advances in Neural Information Processing Systems, 36. arXiv:2305.10601
- Lightman, H., Kosaraju, V., Burda, Y., Edwards, H., Baker, B., Lee, T., Leike, J., Schulman, J., Sutskever, I., & Cobbe, K. (2024). Let's Verify Step by Step. International Conference on Learning Representations (ICLR 2024). arXiv:2305.20050
- Huang, J., Chen, X., Mishra, S., Zheng, H. S., Yu, A. W., Song, X., & Zhou, D. (2024). Large Language Models Cannot Self-Correct Reasoning Yet. International Conference on Learning Representations (ICLR 2024). arXiv:2310.01798
- Kamoi, R., Zhang, Y., Zhang, N., Han, J., & Zhang, R. (2024). When Can LLMs Actually Correct Their Own Mistakes? A Critical Survey of Self-Correction of LLMs. Transactions of the Association for Computational Linguistics. arXiv:2406.01297
- Turpin, M., Michael, J., Perez, E., & Bowman, S. R. (2023). Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting. Advances in Neural Information Processing Systems, 36. arXiv:2305.04388
- Chen, Y., Benton, J., Radhakrishnan, A., Uesato, J., Denison, C., Schulman, J., Somani, A., Hase, P., Wagner, M., Roger, F., Mikulik, V., Bowman, S. R., Leike, J., Kaplan, J., & Perez, E. (2025). Reasoning Models Don't Always Say What They Think. Anthropic. arXiv:2505.05410
- Schluntz, E., & Zhang, B. (2024, December 19). Building Effective AI Agents. Anthropic. https://www.anthropic.com/research/building-effective-agents
- Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K., & Cao, Y. (2023). ReAct: Synergizing Reasoning and Acting in Language Models. Proceedings of the 11th International Conference on Learning Representations (ICLR 2023).
- Shinn, N., Cassano, F., Berman, E., Gopinath, A., Narasimhan, K., & Yao, S. (2023). Reflexion: Language Agents with Verbal Reinforcement Learning. Advances in Neural Information Processing Systems, 36. arXiv:2303.11366
- Walters, W. H., & Wilder, E. I. (2023). Fabrication and errors in the bibliographic citations generated by ChatGPT. Scientific Reports, 13, Article 14045. https://doi.org/10.1038/s41598-023-41032-5
- Gou, Z., Shao, Z., Gong, Y., Shen, Y., Yang, Y., Duan, N., & Chen, W. (2024). CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing. International Conference on Learning Representations (ICLR 2024). arXiv:2305.11738
- Greshake, K., Abdelnabi, S., Mishra, S., Endres, C., Holz, T., & Fritz, M. (2023). Not What You've Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection. Proceedings of the 16th ACM Workshop on Artificial Intelligence and Security (AISec '23). https://doi.org/10.1145/3605764.3623985