Reflexion, annotated
How to read this page
- Any dotted word explains itself when you hover, tab to, or tap it, and so does every symbol in every equation.
- The step-through runs a tiny Reflexion loop for real: code is tested in your browser, failures produce written lessons, and the lessons steer the next attempt.
Each idea climbs the ladder: everyday picture, tiny example, diagram, the math, why it matters. Retry and verification patterns in code are in the planning lesson.
Abstract
“We propose Reflexion, a novel framework to reinforce language agents not by updating weights, but instead through linguistic feedback.”Shinn et al. (2023), Abstract. Read the original
Everyday picture
A student fails a practice exam. They could simply sit it again and hope. Or they could write themselves a short note (“I lost marks because I never checked units; next time, check units before answering”) and read it before the next attempt. Reflexion gives an agent that habit. After each failed attempt it writes a lesson in plain language, keeps the last few lessons in memory, and reads them before trying again. Nothing about the model's weights changes.
What the paper claims
- Agents improve over repeated attempts at the same task by reflecting in words, with no fine-tuning.
- Feedback can be a simple pass/fail signal, a hand-written rule, or the agent's own evaluation (for example, tests it wrote itself).
- Gains over strong baselines: +22 points on ALFWorld, +20 on HotpotQA and +11 on HumanEval, reaching 91% pass@1 on HumanEval against GPT-4's 80% at the time.
Why it matters today
“Try, check against something real, write down what went wrong, retry” is how many coding agents work. This paper is also a careful study of when self-critique helps and when it misleads, which is the practical question.
1 Introduction · original
“This self-reflective feedback acts as a ‘semantic’ gradient signal by providing the agent with a concrete direction to improve upon.”Shinn et al. (2023), §1
Everyday picture
Classic reinforcement learning improves an agent by adjusting millions of weights using a score (a reward), over thousands of attempts. For a huge language model that is slow and expensive, and a bare score says little: “you failed” does not say why. Reflexion turns the score into a sentence that does say why, and keeps it in the prompt.
Tiny example
A reward of 0 tells an agent only that it failed. The reflection “I searched for the director's name but never checked the release year, so I picked the wrong film; next time, confirm the year first” says which step to change and how. Working out which earlier step deserves the blame for a failure is the credit assignment problem, and the reflection is the model attempting it in language.
Advantages and costs, as the paper lists them
- For: no fine-tuning; richer feedback than a number; an explicit, readable memory; concrete hints for the next attempt.
- Against: it depends on the model evaluating itself well, and there is no guarantee it converges.
2 Related work · original
The paper contrasts Reflexion with self-refinement methods that polish a single output, with search methods over actions, and, for code, with systems that debug using tests (AlphaCode, CodeT, Self-Debugging, CodeRL). Its distinguishing ingredients are a persistent memory of reflections across attempts and the ability to work from a binary reward, including in multi-step decision-making tasks. For code specifically, it generates its own tests, so it can report honest pass@1 scores without peeking at the hidden test cases.
3 The method · original
Hover or tap a block, starting with the Actor.
Reading it: within one attempt (a trial), the Actor and the environment exchange actions and observations, building a trajectory. When the trial ends, the Evaluator scores it. If it failed, the Self-Reflection model reads the trajectory and the score and writes a lesson, which is appended to long-term memory. The long top arrow is the only thing that carries over to the next trial: the Actor starts fresh except that it now reads the recent lessons. Every box except the environment can be the same language model with different instructions.
Actor, Evaluator, Self-Reflection, Memory · original
| Part | Everyday picture | In the paper |
|---|---|---|
| Actor | The student sitting the exam | A language model prompted as chain of thought or ReAct, which also reads the memory |
| Evaluator | The marker | Exact-match grading (HotpotQA), a hand-written rule or a model's yes/no judgement (ALFWorld), or self-written unit tests (code) |
| Self-Reflection | The student's note after marking | A language model that turns the trajectory and the score into specific advice |
| Memory | The notebook of notes | Short-term: this trial's trajectory. Long-term: the last Ω reflections, usually Ω = 1 to 3, to fit the context window |
The two kinds of memory map directly onto the terms used for agents today: short-term memory is the working transcript, long-term memory is distilled experience carried between sessions, here a kind of episodic memory. See the memory lesson.
The loop, formally · original
Everyday picture
In ordinary learning, what gets updated between attempts is the student's brain (the weights). In Reflexion, the brain stays fixed and what gets updated is the notebook. The paper's framing is that the agent's “policy”, its way of choosing actions, is the model plus its notebook.
In words: “in each trial, the actor (a fixed model reading its memory) produces a trajectory; the evaluator scores it; the reflector writes a lesson from the trajectory and score; the lesson is appended to memory, which keeps only the most recent few; repeat until the evaluator says pass or trials run out.”
With the numbers: in the step-through below with Ω = 2, trial 1 fails (r₁ = 0), producing sr₁; memory = [sr₁]. Trial 2 fails (r₂ = 0), producing sr₂; memory = [sr₁, sr₂]. Trial 3 reads both lessons and passes (r₃ = 1).
In Python:
# Ω: how many lessons memory keeps
Omega = 2
mem = []
# trials 1 and 2 fail (r_t = 0)
for t, r_t in [(1, 0), (2, 0)]:
# the reflector's lesson from trial t
sr_t = f"sr{t}"
# mem ← mem + [sr_t], keep the last Ω
mem = (mem + [sr_t])[-Omega:]
# what the actor reads in trial 3
mem # → ['sr1', 'sr2']
Why reflection beats plain retrying
If every attempt were an independent coin flip with success probability p, the chance of at least one success in T attempts would be:
In words: “the chance of failing every time is (1 − p) multiplied by itself T times; everything else is at least one success.”
With the numbers: with p = 0.6: one try 0.6, two tries 1 − 0.4² = 0.84, three tries 1 − 0.4³ = 0.936. But this assumes the failures are random. The paper found that baselines retrying at temperature 0.7 on HotpotQA solved none of their first-trial failures in later trials: the failures were systematic, not bad luck. Reflection is what changes p between trials.
In Python:
p = 0.6
# P(success within T)
[round(1 - (1 - p) ** T, 3) for T in (1, 2, 3)] # → [0.6, 0.84, 0.936]
Try it: trial, fail, reflect, retry
The task: write is_palindrome(s), true if s reads the same backwards, ignoring case, spaces and punctuation. The code attempts and reflections are illustrative, written for this page. The tests really run in your browser against a JavaScript version of each attempt.
Long-term memory (what the Actor reads)
Reading it: each press runs one trial: the Actor proposes code, the Evaluator runs the self-written tests (✓ or ✗ each), and on failure the Self-Reflection step writes a lesson into memory. With reflection, each lesson removes one mistake and trial 3 passes every test. With plain retry, nothing carries over, and the same mistake repeats. With weak tests (only two easy cases), trial 1 passes its own tests and is submitted, but fails the hidden tests: a false positive, the failure mode the paper analyses for code in §4.3.
4 Experiments · original
4.1 Sequential decision-making: ALFWorld · original
Everyday picture
The household text game from the ReAct companion: find a spatula in a drawer, chill a tomato in the fridge. The environment only says “done” or nothing, so the agent must judge its own failure. The paper's simple rule: if it repeats the same action with the same result more than 3 times, or takes more than 30 actions, the trial counts as failed and triggers a reflection.
What happened
- ReAct + Reflexion solved 130 of 134 tasks using that simple rule, improving over 12 consecutive trials.
- ReAct alone stopped improving between trials 6 and 7, and stayed stuck with a hallucination rate of 22%.
- A common baseline failure: the agent believes it is holding an item it never picked up, then acts on that belief for many steps. Reflection lets it spot the early mistake in a long, failed trajectory.
4.2 Reasoning: HotpotQA · original
What happened
- Evaluation used exact-match grading as a pass/fail signal between trials, with a memory of 3 reflections, on 100 questions.
- Reflexion improved on both chain-of-thought and ReAct agents across trials; the baselines never solved a question they had failed on the first trial.
- Even when given the ground-truth supporting text (CoT with ground truth), the agent got 39% of questions wrong; reflection improved its accuracy by 14 points.
- Ablation: adding only the previous trajectory as memory helped, but adding a written reflection helped 8 points more. The words matter, not just the replay.
4.3 Programming · original
Everyday picture
For code, the agent writes its own unit tests (up to 6, filtered to those that are syntactically valid), runs its code against them, and reflects on the failures. It never sees the benchmark's hidden tests, so its score is honest “pass@1”.
| Benchmark + language | Previous state of the art | GPT-4 (at the time) | Reflexion |
|---|---|---|---|
| HumanEval (Python) | 65.8 (CodeT + GPT-3.5) | 80.1 | 91.0 |
| HumanEval (Rust) | – | 60.0 | 68.0 |
| MBPP (Python) | 67.7 (CodeT + Codex) | 80.1 | 77.1 |
| MBPP (Rust) | – | 70.9 | 75.4 |
| Leetcode Hard (Python) | – | 7.5 | 15.0 |
Reading it: Reflexion beats the plain model on every benchmark but one, MBPP Python, where it is slightly worse. The next table explains why.
| Benchmark | Base | Reflexion | TP | FN | FP | TN |
|---|---|---|---|---|---|---|
| HumanEval (Python) | 0.80 | 0.91 | 0.99 | 0.40 | 0.01 | 0.60 |
| MBPP (Python) | 0.80 | 0.77 | 0.84 | 0.59 | 0.16 | 0.41 |
| HumanEval (Rust) | 0.60 | 0.68 | 0.87 | 0.37 | 0.13 | 0.63 |
| MBPP (Rust) | 0.71 | 0.75 | 0.84 | 0.51 | 0.16 | 0.49 |
TP: the self-written tests pass and the solution is correct. FN: the tests fail although the solution is correct. FP: the tests pass although the solution is wrong. TN: the tests fail and the solution is wrong.
Reading it: the FP column is the danger. A false positive means the agent's own tests all passed on wrong code, so it stopped early and submitted a bad answer. On HumanEval Python that happened 1% of the time; on MBPP Python, 16%, which is enough to sink its score below the plain model's. False negatives are less harmful: a reflection can decide a test itself is wrong and keep the code. An agent that checks itself is only as good as its checks.
Ablation: which part matters?
| Approach | Tests | Reflection | Pass@1 |
|---|---|---|---|
| Base model | no | no | 0.60 |
| Without test generation | no | yes | 0.52 |
| Without self-reflection | yes | no | 0.60 |
| Reflexion | yes | yes | 0.68 |
Reading it: reflecting without tests is worse than doing nothing (0.52 against 0.60): with no external signal the agent cannot tell whether its code works, keeps editing, and damages good code. Tests without reflection add nothing either: the errors are caught but the fixes do not follow. Only the combination helps. That is the most transferable finding in the paper.
5–7 Limitations, impact and conclusion · original
- Local optima. Optimising through language can still get stuck.
- Small memory. A sliding window of 1 to 3 reflections; the authors suggest vector or SQL databases for more.
- Tests are hard to write for non-deterministic, side-effecting, hardware-dependent or concurrent code.
- Safety. Autonomous agents amplify misuse risks. On the other hand, written reflections make an agent's intentions easier to inspect than the weights of a conventional reinforcement-learning agent.
- Reproducibility note: the authors advise running generated code only in isolated environments.
Using it today
| Finding | What careful systems do | Learn it |
|---|---|---|
| Self-critique without an external signal can make things worse | Pair reflection with external verification: tests, schema checks, database queries | planning |
| Self-written tests produce false positives | Treat agent-written tests as hints; gate releases on human-owned tests and evals | evals |
| Blind retries repeat systematic errors | Carry a short written lesson into the retry, and cap the number of trials | agent loop |
| Memory must stay small | Keep a few distilled lessons, not whole transcripts | memory |
| Generated code is untrusted | Run it in a sandbox with no credentials | guardrails |
Glossary
Every term with hover guidance on this page, in one place.