SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering, annotated
How to read this page
- Any dotted word explains itself when you hover it, tab to it, or tap it, and so does every symbol in every equation.
- The file viewer and the editor are working miniatures of SWE-agent's: scroll a toy file through a window, then try four edits and watch the linter accept or reject them.
- The context chart shows how collapsing old observations keeps a long run inside the context window.
- The trajectory stepper replays one real run from the paper, action by action.
Each idea climbs the ladder: everyday picture, tiny example, diagram, the math, why it matters. The coding agents lesson builds a small coding agent with tools shaped by the same ideas, and the SWE-bench companion explains the benchmark every number here comes from.
Abstract
“We posit that LM agents represent a new category of end users with their own needs and abilities, and would benefit from specially-built interfaces to the software they use.”Yang et al. (2024), Abstract. Read the original
Everyday picture
A carpenter with a good workbench, clamps and a sharp saw builds more than the same carpenter with a kitchen knife. Nothing about the carpenter changed. This paper keeps the model fixed and changes the workbench.
What the paper claims
- An agent-computer interface (ACI): the set of commands an agent can issue and the format of what comes back, designed for a language model the way an editor is designed for a person.
- SWE-agent: GPT-4 Turbo with such an interface resolves 12.47% of the 2,294 SWE-bench tasks, against 3.79% for the best non-interactive system before it.
- Interface design matters on its own: on the 300-task Lite subset, the same model resolves 18.0% with the interface and 11.0% with a plain Linux shell.
- 87.7% on HumanEvalFix, a set of short debugging problems.
Why it matters today
The paper named something every agent builder now thinks about: the tools are part of the system, and how their output reads to a model matters as much as what they can do. The tools lesson teaches the same lesson for any agent.
1 Introduction · original
Everyday picture
People don't write software in a bare terminal if they can help it. They use an editor that shows line numbers, jumps to definitions, underlines typos as they type. A language model given only a Linux shell is a programmer without any of that: every file view is a cat that floods the screen, every edit a fragile sed command, and a mistake produces no warning at all.
Tiny example: the same model, three setups
| Setup | Resolved |
|---|---|
| Retrieval, one prompt, one patch (the SWE-bench paper's baseline) | 2.67% |
| A Linux shell, no demonstration | 7.33% |
| A Linux shell, with a demonstration run | 11.00% |
| SWE-agent's interface | 18.00% |
Reading it: every row is the same model on the same 300 tasks. The first jump, from one-shot retrieval to any interactive shell, is what acting and looking buys: the model can search, run code and try again. The last jump, from shell to SWE-agent, is what better tools buy. The model's weights never change.
The whole paper in one picture
Hover or tap a box. The loop runs clockwise: model, commands, computer, feedback, model.
Reading it: follow the arrows clockwise. The model never touches the computer. It writes a thought and one command, the interface (the dashed frame) runs the command against the repository in a sandboxed shell, and what returns is not the raw output but a view shaped for the model: a numbered window of a file, a short list of search hits, a syntax error with the lines around it. The paper's claim is that both halves of the frame, what the model can say and what it hears back, are design decisions that change results.
Why it matters
Before this paper, work on coding agents mostly asked which model, which prompt, or whether to let it run code. SWE-agent adds a fourth question with a measurable answer: what should the model's tools look like? The paper's two contributions are the concept (the ACI) and an open-source system built on it.
2 The agent-computer interface · original
“The ACI specifies both the commands available to the LM and how the environment state is communicated back to the LM.”Yang et al. (2024), §2
Everyday picture
A car's controls are designed for human hands and eyes: a wheel, pedals, a speedometer you glance at. A self-driving system doesn't need the wheel; it needs precise sensor readings and precise commands. Building one by sitting a robot in the driver's seat to turn the wheel would work badly. Giving a language model a terminal built for people is the same mistake.
Tiny example: what's different about a model as a user
- It cannot skim. A person scrolls past 500 irrelevant lines of
grepoutput in a second. For a model, every line is tokens in its context, paid for and able to distract it. - It cannot see the screen. Graphical editors signal a lot visually. A model needs the same signals as text: line numbers, how many lines are above and below, an explicit error message.
- It is bad at bookkeeping. Working out the arguments for “show me the next 100 lines” from memory, or a multi-line
sedcommand, is exactly the arithmetic models get wrong.
Diagram: four principles
Hover or tap a principle to see the paper's example and the part of the interface that follows it.
Reading it: the top row is about what the model can say: a few simple commands (1), each doing a whole operation in one step (2). The bottom row is about what it hears back: enough to know what happened and no more (3), and a check that catches a mistake before it spreads (4). Every command in §3 can be traced to one or more of these cards.
The math: why compact actions matter
Principle 2 has a simple arithmetic behind it. If each step of a multi-step procedure goes right with probability p, independently, the whole procedure goes right only if every step does.
In words: “the chance that an n-step procedure succeeds is the per-step chance multiplied by itself n times.”
With the numbers (illustrative): a model that gets 90% of commands right edits a three-line block with one edit command 90% of the time. Doing the same with four separate sed calls, each needing the right line number and escaping, succeeds 0.94 = 66% of the time.
In Python:
p = 0.9
# one compact command, then four fiddly ones
round(p ** 1, 2), round(p ** 4, 2) # → (0.9, 0.66)
The chain_success function in the tools lesson is this formula, applied to choosing between five small tools and one larger one.
Why it matters
The authors arrived at these principles the way interface designers do: they watched the agent work on development tasks, noted where it stumbled, changed the interface, and ran a small hyperparameter sweep. The model stays fixed throughout: every gain in this paper comes from the interface.
3 SWE-agent's interface · original
At each step the model writes a thought and a command, in the pattern of ReAct (the ReAct companion takes it apart), and gets the command's result back. Four components make up the interface: search and navigation, a file viewer, a file editor, and context management. Ordinary Linux commands remain available alongside them.
Search and navigation · original
Everyday picture
Ask a librarian where books on Roman roads are and a good one says “three shelves, mostly in history, aisle 4”. A bad one reads you every catalogue entry containing the word “road”.
Tiny example: three commands and one rule
find_file looks for a file by name, search_dir for a string in every file under a directory, search_file for a string in one file. The rule: at most 50 results. Here is search_dir "notes" from one of the paper's printed runs (pylint, Appendix D), shortened:
Found 24 matches for "notes" in /pylint-dev__pylint:
/pylint-dev__pylint/ChangeLog (2 matches)
/pylint-dev__pylint/examples/pylintrc (2 matches)
/pylint-dev__pylint/pylint/checkers/misc.py (9 matches)
...
End of matches for "notes" in /pylint-dev__pylint
One line per file, with a count, and the agent's next thought was that misc.py, with the most matches, was the place to look. When a search is too broad, the reply is a request instead of a flood; from the sympy run: “More than 182 files matched for "Derivative" in /sympy__sympy. Please narrow your search.”
Hover or tap a column.
Reading it: the columns are three ways to search, each tried with GPT-4 Turbo on SWE-bench Lite. The middle one is modelled on editors people like: step through hits one at a time. It did worst, worse even than no search tool at all (12.0% against 15.7%), because given many results the agent dutifully called next until it had visited every one, burning its budget and context. The green column wins by giving the whole picture in a few lines and refusing to give more. (The paper's text calls this comparison Figure 6; it is its Figure 5.)
Why it matters
A search tool for a model should return locations, not contents, and cap its output. The Workspace.search tool in the coding agents lesson does both, with a cap of 20 hits and a request for a narrower pattern above it. The paper notes the cost of the rule: a rejected search means another model call.
The file viewer · original
“The file viewer presents a window of at most 100 lines of the file at a time.”Yang et al. (2024), §3
Everyday picture
Reading a long contract through a letterbox-sized window you can slide up and down, with page numbers printed in the margin and a note saying “12 pages above, 40 below”. You never lose your place, and you never have the whole contract dumped on your desk.
Tiny example: a window on a file
open path shows the first window; goto n jumps so line n is in view; scroll_down and scroll_up move a whole window. Every view starts with the file's path and length and says how many lines are hidden above and below, and every line carries its number. Those numbers are what the edit command takes, so the agent never has to count. (The paper's command table, its Table 4, swaps two descriptions: it says scroll_down moves the window up and scroll_up down.)
Try it: a 30-line toy file, loosely modelled on the pvlib file in the paper's Figure 11. Change the window size, jump to a line, scroll, and read the output exactly as the agent would. Notice what happens to the “more lines” notes, and how many scrolls it takes to see everything with a small window.
Reading it: the grey box is the whole observation the model receives. The first line says which file and how long; the notes in brackets say how much is hidden; each code line starts with its number. With W = 10 the file takes 3 windows to read, with W = 3 it takes 10. At W = 30 the whole file fits, and on a real 1,890-line file that would be the “full file” setting that the paper found hurt.
The math: windows and tokens
In words: “reading a whole file takes the file length divided by the window size, rounded up, views; each view costs the lines it shows times the tokens per line.”
With the numbers: the pvlib file in the paper's Figure 11 has 1,890 lines. At the default W = 100 that is 19 views to read it all, each about 1,000 tokens at an illustrative 10 tokens per line. The whole file in one view would be 18,900 tokens; a 30-line window needs 63 views.
In Python:
import math
L, tau = 1890, 10
# n_views for the default window, then for 30 lines
math.ceil(L / 100), math.ceil(L / 30) # → (19, 63)
# T_view for a 100-line window, and for the whole file at once
min(100, L) * tau, min(L, L) * tau # → (1000, 18900)
Neither extreme wins. In the paper's ablation (§5.1), 30 lines resolved 14.3% of Lite, the whole file 12.7%, and 100 lines 18.0%. Too small a window means many steps and lost context between them; too large a window means the relevant lines are buried.
Why it matters
The viewer turns a stateless shell into a stateful tool: the interface remembers which file is open and where, so “the next 100 lines” is one word, scroll_down, instead of a command whose arguments the model must compute. The Workspace.read_file tool in the coding agents lesson returns a whole file, which is fine for its 7-line files; real repositories need a window.
The editor and its linter · original
“Invalid edits are discarded, and the agent is asked to try editing the file again.”Yang et al. (2024), §3
Everyday picture
A word processor that refuses to save a document with a sentence that doesn't parse, shows you the sentence as you typed it next to the version before, and tells you what's wrong. Annoying for a person on a good day; invaluable for someone who can't see the page.
Tiny example: edit n:m
edit 404:407 followed by some lines and end_of_edit replaces lines 404 to 407 of the open file with those lines, in one step. After the edit, a linter (flake8, limited to a few codes for undefined names, bad indentation and syntax errors) checks the file. If it complains, the edit is reverted and the agent gets a message with three parts: the errors, how the file would have looked, and how it looks now. If not, the viewer shows the updated lines at once.
Try it: the task is to remove the parameter orientation_strategy from basic_chain in the toy file. Pick an edit and see what the interface says. The checker here is a tiny stand-in for flake8 that knows three of its codes: F821 (undefined name), E111 (indentation not a multiple of four) and E999 (syntax error).
Reading it: the grey box is the agent's next observation. Deleting the parameter first (the paper's own Figure 11 example) leaves two uses of the name in the body, so F821 fires twice and nothing changes. Deleting lines 7 to 11 in one go takes line 8 with it, the one that closes the signature's bracket: a syntax error. Re-indenting with three spaces trips E111. Only the order the guardrail forces works: remove the code that uses the parameter first, then the parameter. The paper names this trade-off: the guardrail sometimes dictates the order of edits.
The math: how often editing goes wrong, and recovers
From the paper's analysis of GPT-4 Turbo on the full benchmark: 1,185 of 2,294 runs (51.7%) had at least one edit rejected by the linter. What happens next can be measured.
In words: “the recovery rate is the number of runs of failed edits that end in a successful edit, divided by all runs of failed edits.”
With the numbers: in the 286 resolved runs, 135 streaks of failed edits, 104 recovered: 0.77. Across all runs, 810 recovered and 555 did not: 0.59. The paper also gives the chance an edit attempt eventually succeeds: 90.5% for a fresh attempt, 57.2% once one failure has happened, so 42.8% of the time a first failure is never recovered from.
In Python:
# r_rec = S / (S + U), in resolved runs, then in all runs
S, U = 104, 135 - 104
round(S / (S + U), 2) # → 0.77
round(810 / (810 + 555), 2) # → 0.59
# a first failure that is never recovered from
round(1 - 0.572, 3) # → 0.428
# runs with at least one failed edit, of the 2,294
round(100 * 1185 / 2294, 1) # → 51.7
One number in this appendix doesn't add up: the paper says 113 of the 286 resolved runs, “31.5%”, had a failed edit, but 113 / 286 is 39.5%.
round(100 * 113 / 286, 1) # → 39.5
Why it matters
A bad edit that is kept breaks the file, the next run of the code fails for a reason unrelated to the issue, and the agent chases its own mistake: the paper's “cascading” failures. The linter stops that at the first step. Turning linting off cost 3 points on Lite (18.0% to 15.0%), and without the edit command at all (only sed and shell redirection) the agent fell to 10.3%. The Workspace.edit_file tool in the coding agents lesson is built on the same idea from the other side: it refuses an edit whose target text is missing or ambiguous, and says how to fix the call.
Context management · original
Everyday picture
A detective's notebook keeps the last few pages of notes in full and, for everything older, one line each: “Tuesday: interviewed the gardener, alibi checks out.” The plan and the trail survive; the transcripts don't clutter the desk.
Tiny example
SWE-agent keeps the system prompt, the demonstration and the issue at the top of every call. Of the observations (the tool outputs), only the last 5 are shown in full; each older one is collapsed to a single line. Two more rules: a command with no output gets an explicit line (“Your command ran successfully and did not produce any output”) so the agent doesn't run extra commands to find out what happened; and when a malformed reply is followed by a valid one, the malformed reply and its error message are removed from the history.
The math: how big the context gets
In words: “with full history, the context after t steps is the fixed start plus every observation so far; with the last-5 rule, it is the fixed start, the five newest observations in full, and one short line for each older one.” (For t of 5 or fewer the two are the same.)
With the numbers (illustrative): a fixed start of 2,000 tokens, observations of 800 tokens, collapsed lines of 10 tokens. After 20 steps, full history holds 2,000 + 20 × 800 = 18,000 tokens; the last-5 rule holds 2,000 + 5 × 800 + 15 × 10 = 6,150. The model's own thoughts and commands are left out of both for simplicity.
In Python:
S, o, c, t = 2000, 800, 10, 20
obs = [o] * t
# C_full(t): every observation in full
S + sum(obs) # → 18000
# C_5(t): the last five in full, one line for each older one
S + sum(obs[-5:]) + (t - 5) * c # → 6150
Hover the chart, or tab to it and use the arrow keys, to read the context size at each step.
Reading it: the x-axis is the step of the run, up to 40 (the paper's $4 budget ended runs somewhere between steps 30 and 40); the y-axis is tokens in the context, from the formula above with a 2,000-token start and 10-token collapsed lines (illustrative). The full-history line climbs by a whole observation per step; the last-5 line flattens after step 5 and climbs by one short line per step. The flat dotted line is GPT-4 Turbo's 128,000-token window. Push the slider to 3,000 and full history nearly fills the window by step 40, while last-5 stays under 20,000.
Why it matters
In the paper's ablation, keeping the full history cost 3 points on Lite (18.0% to 15.0%), and the hyperparameter sweep in Appendix B preferred last-5 for GPT-4 Turbo at both window sizes and both temperatures. Collapsing isn't only about fitting: stale file views in the history can mislead the agent about what a file now contains. The context lesson builds the same idea as a rolling summary in summarize_turns, and compress_tool_output trims each tool result to what the task needs.
4 Experimental setup · original
Everyday picture
To show a new workbench helps, give the same carpenter the same jobs with the old tools and the new, and count finished jobs. Also count the cost, since the new workbench might just be slower and more thorough.
Tiny example: the setup in five lines
- Tasks: all 2,294 of SWE-bench for the headline, the 300 of SWE-bench Lite for ablations, and HumanEvalFix (164 short buggy functions per language) for basic editing.
- Models: GPT-4 Turbo (128k-token window) and Claude 3 Opus (200k). Llama 3 and DeepSeek Coder were tried and did poorly as agents; Llama 3's 8k window was too small.
- Baselines: the SWE-bench paper's retrieval setup (BM25 files in one prompt, one patch out), and a Shell-only agent in the InterCode environment, which gets a Linux shell and nothing else.
- Metrics: % resolved (pass@1, one run per task) and the average API cost per resolved task.
- Budget: $4 per task; a run that hits it has its edits so far submitted automatically.
Reading it: each system has two bars, from Table 1 of the paper: the percentage of SWE-bench Lite resolved (solid, axis 0 to 20%) and the average cost (striped, axis 0 to 3 dollars). Read down the solid bars: retrieval at the top of the list, the plain shell in the middle, SWE-agent at the bottom for both models, with the longest bars. The striped bars go the other way: acting costs more than one prompt. On the full 2,294 tasks the paper reports SWE-agent at 12.47% with GPT-4 Turbo and 10.46% with Claude 3 Opus.
The math: points, ratios and folds
The paper reports its gains three ways, and they sound very different.
In words: “the difference in percentage points subtracts the rates; the relative gain divides that difference by the old rate; the fold is the new rate divided by the old.”
With the numbers: SWE-agent against the shell with a demonstration, 18.00% against 11.00%: 7.0 points, a 64% relative gain (the paper's “64%”). Against the shell without a demonstration, 7.33%: 10.67 points (the introduction's “10.7 percentage points”, which compares with that weaker shell). Against retrieval, 2.67%: 6.7 times (the paper's “6.7-fold”). And the cost, in dollars per resolved task: 1.67 against 0.13 is 12.8 times, and for Claude 3 Opus 2.18 against 0.25 is 8.7 times, the paper's “8-13x”.
In Python:
a = 18.00
# b: the shell with a demonstration
b = 11.00
round(a - b, 2), round((a - b) / b, 3) # → (7.0, 0.636)
# b: the shell without one, then retrieval
round(a - 7.33, 2), round(a / 2.67, 1) # → (10.67, 6.7)
# cost folds, GPT-4 Turbo then Claude 3 Opus
round(1.67 / 0.13, 1), round(2.18 / 0.25, 1) # → (12.8, 8.7)
HumanEvalFix
| Model | Python | JS | Java |
|---|---|---|---|
| CodeLLaMa-instruct-13B | 29.2 | 19.5 | 32.3 |
| GPT-4 | 47.0 | 48.2 | 50.0 |
| DeepseekCoder-CodeAlpaca-6.7B | 49.4 | 51.8 | 45.1 |
| WaveCoder-DS-6.7B | 57.9 | 52.4 | 57.3 |
| SWE-agent with GPT-4 Turbo | 87.7 | 89.7 | 87.9 |
Reading it: HumanEvalFix gives a buggy function and its tests and asks for the fix, with no repository to search. SWE-agent runs with the same configuration apart from a demonstration in the right language, and beats the best earlier score in each language by about 30 points or more. The paper's §5 gives the result as 88.3%, while its abstract and this table give 87.7% for Python; neither 88.3 nor any average of the three columns (88.4) matches exactly.
Why it matters
A result reported only as “64% better” or only as “7 points better” can mislead in either direction; the page's formula shows why both are true at once. The cost column matters as much: an agent that resolves more tasks but costs 13 times as much per attempt has to be judged on cost per resolved task, which the coding agents lesson computes in cost_per_resolved.
Six runs and pass@k · original
Everyday picture
A basketball player who makes 18% of shots from half court will usually miss, but let them shoot six times and they may well score once. Which number describes them depends on how many shots you allow.
Tiny example
The authors ran SWE-agent with GPT-4 Turbo on Lite six times. The six resolved rates were 17.33, 18.00, 18.00, 18.67, 17.33 and 18.33%: an average of 17.94%, with little spread. But the tasks solved changed between runs: counting a task solved if any of k runs solved it, the rate rises to 32.67% at k = 6.
In words: “for one task run n times and solved in c of them, pass@k is one minus the share of all groups of k runs that contain no success; the benchmark's figure averages this over tasks.” This is the unbiased estimator from the HumanEval companion; the paper does not print its formula, but its pass@1 equals the mean of the six runs, which this estimator guarantees.
With the numbers (illustrative): a task solved in c = 2 of n = 6 runs has pass@1 = 2/6 = 0.33, pass@3 = 1 − C(4, 3)/C(6, 3) = 1 − 4/20 = 0.8, and pass@6 = 1. From the paper's Table 10, the mean of the six runs is 17.94%.
In Python:
from math import comb
n, c = 6, 2
[round(1 - comb(n - c, k) / comb(n, k), 2) for k in (1, 3)] # → [0.33, 0.8]
# pass@1 over the benchmark is the mean of the six runs
runs = [17.33, 18.00, 18.00, 18.67, 17.33, 18.33]
round(sum(runs) / len(runs), 2) # → 17.94
Hover the chart, or tab to it and use the arrow keys, to read pass@k at each k.
Reading it: the x-axis is k, the number of runs allowed; the y-axis is the percentage of Lite tasks solved by at least one of them, from Table 10 of the paper. The curve climbs fast and then flattens: the second run adds 6 points, the sixth only 1.3. The gap between 17.94% and 32.67% is tasks the agent can solve but doesn't reliably: its per-task behaviour varies a lot even though its average barely moves.
Why it matters
A user running an agent once gets pass@1. The gap to pass@6 is the prize for anything that can tell a good run from a bad one, such as running the tests or a verifier, which is why so much later work generates several attempts and picks one. The coding agents lesson computes the same estimator in pass_at_k.
5 Results · original
“An LM-friendly ACI's value is confirmed by SWE-agent's 64% relative increase compared to Shell-only, both with GPT-4 Turbo.”Yang et al. (2024), §5
The headline: 12.47% (286 of 2,294) of the full benchmark and 18.00% (54 of 300) of Lite, both with GPT-4 Turbo. Two appendix results round it out: success is unrelated to the issue's age (a check against memorized fixes), and the interactive agent finds the right files more often than retrieval does, with a file-level F1 score of 59.05% against 45.47% for BM25 with Claude 3 Opus.
# 286 of 2,294 and 54 of 300
round(100 * 286 / 2294, 2), round(100 * 54 / 300, 2) # → (12.47, 18.0)
5.1 What each part of the interface is worth · original
Everyday picture
To find out which of your habits makes you sleep well, change one at a time and keep a diary. An ablation does that for a system: swap one component for an alternative, keep everything else, and measure.
Tiny example: one swap
Replace the 100-line viewer with one that shows whole files, keep everything else: Lite drops from 18.0% to 12.7%, 16 fewer tasks resolved out of 300.
# 18.0% and 12.7% of 300 tasks
round(0.180 * 300), round(0.127 * 300) # → (54, 38)
Reading it: Table 3 of the paper, SWE-bench Lite with GPT-4 Turbo. Bars are grouped by component; the striped bar in each group is SWE-agent's choice, which is always 18.0% because it is the same full system. Every alternative is lower. The biggest drops come from the editor (no edit command: 10.3%) and search (iterative search: 12.0%); the smallest from dropping the demonstration (16.3%). Each number is one run of 300 tasks, and the six repeated runs varied between 17.33% and 18.67%, so differences below a point or so are within run-to-run noise.
Why it matters
Three lessons carry beyond this paper. Interfaces designed for people can be worse than nothing for a model (iterative search below no search). Consolidating an operation into one command helps more than any other single change (the editor). And feedback has a size sweet spot (the viewer). The authors compare their design process to the user studies of human-computer interaction, with a model as the user.
5.2 How the agent behaves · original
Everyday picture
Watch an experienced engineer fix a bug and you see a rhythm: reproduce it, find where it lives, change something, run again, repeat, and stop when it's fixed. The paper finds the same rhythm in the agent's logs, without anyone programming it in beyond a tip in the prompt.
Tiny example: the most common openings
In the 286 resolved runs, the three-action pattern create, edit, python (make a file, write a reproduction script into it, run it) starts at turns 1 to 3 in 156 runs, by far the most common. The alternative opening is localization: search_dir, open, search_file (21 runs). From turn 5 on, the most common actions are edit and python, alternating. A frequent ending is python, rm, submit: run the check once more, delete the scratch script, submit.
Hover or tap a phase.
Reading it: runs begin at the top and go to either reproduction or localization first; the arrow between them shows that one often leads into the other. Then comes the small loop in the middle, edit and run, which the paper's action counts show dominating from turn 5 onward, with more searching and scrolling mixed in when a change breaks something elsewhere. Two exits: an intentional submit (green, thick border), or running out of budget, when the edits so far are submitted automatically.
The math: succeeding quickly, failing slowly
In words: “among runs that ended a given way, the share that resolved their task.”
With the numbers: from Table 13 of the paper (GPT-4 Turbo, full benchmark), 1,589 runs ended with submit and 266 of them resolved their task: 16.7%. 630 ran out of budget with edits to submit, and 20 resolved: 3.2%. Of the 286 resolved runs, 266 (93.0%) submitted on their own. The paper's text gives 14.3% for the first figure, which its own table doesn't reproduce, and its Table 13 rows for all tasks add up to 2,268 rather than 2,294 (and to 2,004 for Claude 3 Opus).
In Python:
# P(resolved | submit) and P(resolved | out of budget)
round(100 * 266 / 1589, 1), round(100 * 20 / 630, 1) # → (16.7, 3.2)
# resolved runs that submitted on their own
round(100 * 266 / 286, 1) # → 93.0
# the "All" row of Table 13 for GPT-4 Turbo
sum([1589, 630, 48, 1]) # → 2268
Claude 3 Opus shows a similar mismatch: Table 1 gives it 13.00% of Lite, which is 39 tasks, while Table 13 and Figure 14 count 35 resolved Lite runs (11.67%).
round(0.13 * 300), round(100 * 35 / 300, 2) # → (39, 11.67)
Resolved runs are short: a median of 12 turns and 1.21 dollars, against 21 turns and 2.52 dollars for unresolved ones (the paper calls the second pair a mean, the first a median). So the authors doubt a bigger budget would help much: if the first 10 to 20 turns don't find a fix, later turns mostly keep editing locally instead of trying a different approach.
Why failures happen
For the 248 unresolved Lite runs, GPT-4o sorted each trajectory into one of nine categories; on 15 runs the authors labelled by hand, it agreed with them 87% of the time. About half the failures (52.0%) are an incorrect or overly specific implementation: a change in a reasonable place that doesn't fix the issue, or fixes only the exact case in the issue. Another 23.4% are failed edit recovery, the agent stuck in a loop of rejected edits. The rest, under 25% together, include not reproducing the bug, never opening the right file, running out of budget, and giving up early.
Why it matters
An agent's log is a dataset. Counting actions per turn, which pairs follow which, and how runs end tells you where the interface or the model falls short, the way a product team reads usage analytics. The paper's findings (fast successes, slow failures, cascading edits) are the kind that change how a system is built: cap budgets, detect edit loops, and invest in better editing.
Try it: one real run, step by step
Everyday picture
A flight recorder, replayed: every control input and every instrument reading, in order.
Try it: step through SWE-agent's successful run on psf__requests-2317 from the paper's Appendix D. Each step shows the command the agent issued, a summary of its reasoning, and what came back (tool output quoted briefly, the rest summarized).
Reading it: the issue says requests turns the method b'GET' into the string "b'GET'" on Python 3, so servers answer 404. Steps 1 to 4 are localization: find the file, open it, search it for the call named in the issue, jump to line 428. Step 5 is the fix, one edit. Steps 6 to 8 check it with a small script. Steps 9 and 10 clean up and submit. Ten actions, no failed edit. The evaluation passes, though the real fix used an existing helper, to_native_string, where the agent wrote its own decode: the same habit of ignoring existing utilities that the SWE-bench paper found in single-shot patches.
Why it matters
This is the edit-run-test loop of the coding agents lesson (fix_until_green, driven there by the scripted careful_fixer) running on a real repository. One difference is worth noticing: here the agent decides when it's done and submits, and the hidden tests judge afterwards. In the lesson's harness, a test run decides “done”, and a model that claims success without passing tests is sent back to work.
6 Related work and 7 Discussion · original
Everyday picture
Human-computer interaction as a field grew up studying how people use software and designing it around them. The authors propose the same discipline for a new kind of user.
What came before
Code benchmarks (HumanEval and its extensions) were short and self-contained, and the top method on HumanEval had reached 94.4%. Software engineering benchmarks such as SWE-bench asked for repository-level work. Language agents had spread to web navigation, computer control and code, usually with interfaces made for people (a shell, a Python interpreter). The authors believe SWE-agent is the first agent for end-to-end software engineering.
Limitations, risks, and what's next
- Small toolkit: no web browsing or static analysis yet, though the configuration makes adding tools easy.
- Manual design: the interface and its prompt tips came from people reading trajectories; automating that is open.
- Scope: software engineering only; whether the principles transfer to other domains is untested.
- Safety: model-written code runs in throwaway Docker containers, for inference and for evaluation, which the authors note are weaker isolation than virtualized hardware but adequate for code not written to escape. They also warn of fake “harnesses” that plant malicious instructions in issue text, and of agents producing malicious code.
Why it matters
The containment point is the same one the coding agents lesson makes with run_sandboxed: in-process limits are useful, but code a model writes belongs in a separate, disposable environment with no network or secrets it doesn't need.
Appendix B: tuning and outcomes · original
Everyday picture
Before a race, a team tries a few tyre pressures on a short practice lap. They don't test every combination on the real course.
Tiny example: the sweep
| Temperature | Window | History | Resolved |
|---|---|---|---|
| 0.0 | 100 | Last 5 observations | 15.1 |
| 0.0 | 100 | Full | 14.1 |
| 0.0 | 200 | Last 5 observations | 10.8 |
| 0.0 | 200 | Full | 9.2 |
| 0.2 | 100 | Last 5 observations | 12.4 |
| 0.2 | 200 | Full | 8.7 |
Reading it: the highlighted row is the configuration used for every main result: temperature 0, a 100-line window, last-5 history. For GPT-4 Turbo, the smaller window beat 200 lines in every pairing, and last-5 beat full history in every pairing. Claude 3 Opus tied three ways at 8.1%, including this configuration, so the same one was used for both models. With only 37 tasks, one task is 2.7 points, so small differences here are fragile.
# one task out of 37, in percentage points
round(100 / 37, 1) # → 2.7
Larger patches than the real fixes
Unlike the one-shot models of the SWE-bench paper, whose patches were smaller than the real fixes, SWE-agent's patches are larger: over all tasks, GPT-4 Turbo's median patch adds 12 lines against 7 for the real fix. The authors put most of the difference down to reproduction code left in the patch. Resolved patches are smaller (a median of 3 lines added), consistent with multi-place changes still being hard.
Why it matters
A tuning set of 37 tasks and an ablation on 300 are what a single research group can afford at a few dollars per run. Reading the results well means keeping their noise in mind, and the paper's own repeated runs give a good yardstick for it.
Appendix C: the prompts · original
“Across templates, we find that providing tips which tell agents to not make specific mistakes, avoid common pitfalls, and use helpful execution signals are effective for eliciting more successful problem solving.”Yang et al. (2024), Appendix C
Everyday picture
A new hire's first-day pack: how the office works, one worked example of a finished job, today's assignment with a list of tips collected from everyone who stumbled before, and then short daily updates.
Tiny example: the messages of a run
Hover or tap a message.
Reading it: the three boxes at the top are sent once and never collapsed. Then the loop: every model reply is checked for the format (one discussion, one command). A malformed reply gets the error message on the left; three malformed replies in a row end the run. A valid command's output goes back through the next-step template, which appends the open file and working directory; a command with no output gets the explicit message on the right. The long arrow closes the loop.
The tips, briefly
The issue message ends with six tips, each written after watching agents fail on development tasks: reproduce the bug first and re-run the reproduction after fixing; if a command fails, don't repeat it unchanged; use goto 583 instead of scrolling repeatedly; search the repository for how an input file is used; watch which file is open and which directory you're in; check every edit for wrong line numbers or indentation. The system prompt adds two warnings in capitals, one about indentation and one against issuing two commands at once. The authors also note that the system prompt says interactive programs like python or vim can't be used, since the environment has no interactive session; scripts can still be run.
Why it matters
Much of an agent's behaviour is set in plain English, and the authors say so: their tips don't scale as a method, but they work. Several tips also have a matching guardrail in the tools, which enforces what the prompt only asks: the linter backs up the reminder about indentation, and the search cap backs up the advice to search specifically. The tools lesson makes the same point about tool descriptions as prompts.
What happened next
| In the paper | What it became | Learn it |
|---|---|---|
| Commands and feedback designed for a model | Tool design as a first-class part of any agent: compact results, explicit errors, few well-described tools | tools lesson |
| Search capped at 50 results | Search that returns locations and asks for a narrower query | Workspace.search |
| Edits checked by a linter before they are kept | Edit tools that fail loudly and say how to retry | Workspace.edit_file |
| Last 5 observations in full | Rolling summaries and trimmed tool output | summarize_turns |
| Docker containers for model-written code | Sandboxes with no network or secrets | run_sandboxed |
| 12.47% on SWE-bench | The benchmark's grading rules, explained | SWE-bench companion |
Glossary
Every term with hover guidance on this page, in one place.