An annotated companion · AI Primer

SWE-bench: Can Language Models Resolve Real-World GitHub Issues?, annotated

About this page. This is a companion, not a copy. It follows the paper section by section, quotes only a sentence or two per section (clearly marked), and explains everything in its own words. The paper is published on arXiv under a Creative Commons Attribution 4.0 licence (CC BY 4.0), which allows reproducing its tables with credit: the numbers restated here from its tables are attributed where they appear, and its figures are redrawn from scratch rather than copied. It follows the latest arXiv version (v3, 2024), which adds results for later models and the 300-task Lite subset. Equations are this page's own way of writing the paper's grading rules, with every symbol decoded. Numbers the paper does not give are labelled illustrative, and places where the paper disagrees with itself are pointed out where they occur.

How to read this page

  • Any dotted word explains itself when you hover it, tab to it, or tap it, and so does every symbol in every equation.
  • The context budget lets you choose how many tokens of retrieved code a model may read, and watch recall and clutter rise together.
  • The grader runs six patches for one small bug through the benchmark's real grading rule and places each in the paper's grid of outcomes.

Each idea climbs the ladder: everyday picture, tiny example, diagram, the math, why it matters. The coding agents lesson builds a three-task SWE-bench of its own, graded the same way, in plain Python.

Abstract

“Given a codebase along with a description of an issue to be resolved, a language model is tasked with editing the codebase to address the issue.”Jimenez et al. (2023), Abstract. Read the original

Everyday picture

A school exam asks for a short essay on a fresh page. A job asks you to fix a problem in a thousand-page report that twenty other people wrote, without breaking anything they rely on. Earlier code benchmarks were the exam. SWE-bench is the job.

What the paper claims

  • A benchmark of 2,294 tasks, each a real GitHub issue from one of 12 popular Python repositories, paired with the pull request that resolved it.
  • The task: given the issue text and the whole repository, produce an edit that makes the project's own tests pass.
  • The finding: at the time, even the best model, Claude 2, resolved only 1.96% of the tasks, when the files to read were picked by a simple search engine.
  • A training set of about 19,000 further issues from 37 other repositories, and two fine-tuned models, SWE-Llama 7b and 13b.

Why it matters today

SWE-bench turned “can it write a function?” into “can it do a software engineer's task?”, and graded it by running tests rather than by comparing text. It became a standard measure for coding agents, starting with the SWE-agent companion's paper, which lifted the resolved rate to 12.5% by letting a model work inside the repository.

1 Introduction · original

“Coding tasks are appealing as they pose challenging problems to LMs yet generated solutions can be easily verified by running unit tests.”Jimenez et al. (2023), §1

Everyday picture

A good exam has two properties that pull against each other: it must be hard enough that the best students still miss questions, and each answer must be quick to mark. Code gives both. The HumanEval companion shows the first generation: short, self-contained functions checked by unit tests. By 2023 models were scoring well on those. Real software work is different: find the right lines among thousands of files, understand how they interact, change several places at once.

Tiny example: an exercise against a task

A HumanEval problem against the average SWE-bench task. HumanEval figures from Chen et al. (2021); SWE-bench figures from Table 1 of Jimenez et al. (2023).
HumanEval problemSWE-bench task (average)
What the model is givena function signature and docstringan issue of 195 words and a repository of 3,010 files (438,000 lines, tests not counted)
What it must produceone function bodyan edit touching 1.7 files, 3 functions and 32.8 lines
How it is graded7.7 hidden unit teststhe repository's own tests: 9.1 that must start passing, 120.8 run in all

Reading it: the grading idea is the same in both columns: run the code, trust the tests. What changes is everything around it. The input grows from a few lines to hundreds of thousands, the output from one function to changes in several places, and the tests from a handful written for the benchmark to the project's real test suite.

The whole paper in one picture

Issue text195 words on average Codebase snapshot3,010 files on average Language model Generated patch Evaluation in a fresh copy Apply patch andthe PR's hidden tests,then run them Resolved only ifevery fail-to-passand pass-to-passtest passes

Hover or tap a box. The model's work ends at the patch; everything inside the dashed frame is the benchmark's.

Figure 1 of the paper, redrawn and expanded with the grading rule of §2.2 and Appendix A.4. Based on Jimenez et al. (2023), Figure 1.

Reading it: start at the top. An issue and a frozen copy of the repository, as it stood just before the fix was written, go to the model. The model writes a patch: a list of lines to remove and add. Only the patch crosses into the dashed frame, where a fresh copy of the repository receives it along with the tests that the real pull request added. The green box with the thick border is the whole verdict: the tests that the real fix made pass must now pass, and every test that already passed must still pass.

Why it matters

Three things make this a durable design. The tasks are real, so a score means something about real work. Grading is by execution, so there is no human in the loop and no argument about style. And collection is automatic, so new tasks can be gathered from issues filed after a model was trained, which the authors offer as a defence against contamination.

2 SWE-bench · original

Four parts: how 93,139 pull requests were filtered down to 2,294 tasks, what exactly the model must do and how it is graded, what makes the tasks hard, and a smaller subset for cheaper experiments.

2.1 Building the benchmark · original

“We filter out task instances without at least one test where its status changes from a fail to pass (henceforth referred to as fail-to-pass test).”Jimenez et al. (2023), §2.1

Everyday picture

A teacher building an exam from last year's homework keeps only the questions that had a marked answer key, and only those where the key actually distinguishes right from wrong. A pull request is a homework answer; its new tests are the answer key. SWE-bench keeps a pull request only if its tests fail on the old code and pass on the new.

Tiny example: one pull request through the three stages

Take a hypothetical pull request to a date library titled “Fix leap years for century years, fixes #412”. Stage I: it comes from a popular repository, so it is collected. Stage II: it is merged, it names an issue it resolves (“fixes #412”), and it edits a file whose path contains test, so it becomes a candidate. Stage III: in a clean environment, the new tests run on the old code, is_leap(1900) fails; with the fix applied, it passes. At least one test went from fail to pass, the install worked, so it becomes a task. A pull request that only fixed a typo in the docs would have left at Stage II; one whose tests passed both before and after would have left at Stage III.

Stage I: scrape pull requests12 popular repositories: 93,139 PRs Stage II: attributesmerged, resolves an issue, edits tests11,407 candidates Stage III: executioninstalls, runs, and at leastone fail-to-pass test2,294 tasks

Hover or tap a stage to see what it checks and what it removes.

Figure 2 of the paper, redrawn as a funnel, with the counts of its Table 10. Based on Jimenez et al. (2023), Figure 2 and Table 10.

Reading it: each box is narrower because fewer pull requests survive it. The first two stages read only metadata (is it merged, does it mention an issue, did it touch a test file), so they are cheap. The third actually installs the repository at that moment in its history and runs the tests twice, once before and once after the fix, which is expensive but is the only way to know the tests really check the fix. The paper's text rounds the first count to about 90,000; its Table 10 gives 93,139.

The math: how much survives

In words: “the share of pull requests that become tasks is the share that pass the attribute filter times the share of those that pass the execution filter.”

With the numbers: 11,407 / 93,139 = 0.122 pass Stage II; 2,294 / 11,407 = 0.201 of those pass Stage III; together 0.0246, about one pull request in 41. By the paper's account, install and run failures alone remove about half of the candidates in Stage III; the fail-to-pass check and the log checks of Appendix A.3 remove most of the rest. Flask is the extreme: 2,434 pull requests, 107 candidates, 11 tasks.

In Python:

N_PR, N_cand, N_task = 93139, 11407, 2294
# Stage II keeps this share
round(N_cand / N_PR, 3)  # → 0.122
# Stage III keeps this share of the candidates
round(N_task / N_cand, 3)  # → 0.201
# Y, the whole funnel, and "one in how many"
Y = N_task / N_PR
round(Y, 4), round(1 / Y)  # → (0.0246, 41)
# flask, the smallest repository
round(11 / 2434, 4)  # → 0.0045

Where the tasks come from

Reading it: each bar is one repository's number of final tasks, the paper's Figure 3 redrawn from its Table 10. Django alone supplies 850 of 2,294 (37%) and sympy 386, while flask gives 11 and seaborn 22. Pick a repository to read its whole funnel. The imbalance matters when reading scores: an overall resolved rate is mostly a Django and sympy rate, and a model that is good at one of them moves the total far more than one good at flask.

Why it matters

The fail-to-pass requirement is what makes every task checkable: each has at least one test that the old code fails, so doing nothing cannot score. The same filter has a cost the paper names in Appendix A.3: tasks whose tests call a function the fix newly created are removed, because nothing in the issue says what that function should be called, and no one could be expected to guess it.

2.2 The task and its grade · original

“If the patch applies successfully and all of these tests pass we consider the proposed solution to have successfully resolved the issue.”Jimenez et al. (2023), §2.2

Everyday picture

An editor sending corrections to a printer doesn't resend the whole book. They send a list: “page 12, line 4, delete this sentence, put this one instead”. A patch file is that list for code. The printer applies it mechanically, and if a line it was told to change is not where the list says, the correction bounces.

Tiny example: reading a patch

Here is a patch that fixes the leap-year bug from the coding agents lesson's toy repository. The header names the file, the @@ line says where the change sits, lines starting with - are removed, lines starting with + are added, and lines starting with a space are unchanged context that must match exactly for the patch to apply.

--- a/dates.py
+++ b/dates.py
@@ -1,2 +1,2 @@
 def is_leap(year):
-    return year % 4 == 0
+    return year % 4 == 0 and (year % 100 != 0 or year % 400 == 0)

The model's answer to a task is exactly such a file. The grader applies it with the Unix patch program, adds the hidden tests, and runs them. Suppose the task has two fail-to-pass tests (is_leap(1900) and is_leap(2100) must be False) and four pass-to-pass tests (2000 and 2024 are leap years, 2023 is not, February 2024 has 29 days). The patch above passes all six: resolved. A patch that forgot the 400 rule would pass both fail-to-pass tests and break 2000: not resolved.

The math: the resolved rate

In words: “a task scores 1 only if the patch applies and every one of its fail-to-pass tests and every one of its pass-to-pass tests passes; a single failure anywhere multiplies the whole score to 0. The resolved rate is the share of tasks that score 1.”

With the numbers: the patch that forgets the 400 rule: applies = 1, both fail-to-pass tests pass (1 · 1), but the pass-to-pass product is 0 · 1 · 1 · 1 = 0, so r = 0. Across the benchmark, Claude 2 given the files the real fix edited (the “oracle” setting of §4.1) resolves 110 tasks: R = 110 / 2,294 = 4.80%.

In Python:

import math
def is_leap(year):
    # the patch that forgets the 400 rule
    return year % 4 == 0 and year % 100 != 0
applies = 1
# 1[pass(t)] for each fail-to-pass test, then each pass-to-pass test
F = [is_leap(1900) == False, is_leap(2100) == False]
P = [is_leap(2000) == True, is_leap(2024) == True, is_leap(2023) == False, (29 if is_leap(2024) else 28) == 29]
F, P  # → ([True, True], [False, True, True, True])
# r_i: one failure anywhere zeroes the product
r = applies * math.prod(F) * math.prod(P)
r  # → 0
# R = (1/N) Σ r_i: Claude 2, oracle files, 110 of 2,294 tasks
round(100 * 110 / 2294, 2)  # → 4.8

There is no partial credit, and there is no reading of the patch. A patch that is ugly, or that solves the problem differently from the real fix, counts as long as the tests pass; a patch that is almost right counts for nothing. grade in the coding agents lesson is this rule, returning a Grade with both test reports, and resolved_rate is R.

Why it matters

The pass-to-pass tests are what stop a fix from being a trade. A patch that makes the issue go away by breaking something else is common (Appendix C counts them), and without those tests it would score. The two-list design, fail-to-pass for “the issue is fixed” and pass-to-pass for “nothing else broke”, has become a common way to grade code changes.

2.3 What makes it hard · original

Everyday picture

Finding the one wrong word in a book is easy when someone tells you the page. SWE-bench tells you the symptom, in the words of a user who hit it, and hands you the library.

Tiny example: the needle and the haystack

Selected rows of Table 1 of Jimenez et al. (2023): average and largest values per task, over all 2,294 tasks.
What is measuredMeanMax
Issue length (words)195.14,477
Codebase files (not tests)3,0105,890
Codebase lines (not tests)438K886K
Lines edited by the real fix32.85,888
Files edited by the real fix1.731
Functions edited by the real fix336
Fail-to-pass tests9.11,633
Tests run in all120.89,459

Reading it: the highlighted row against the one two above it is the difficulty in one comparison. The medians in the paper's Appendix A.5 tell the same story for a typical task: an issue of 140 words, a codebase of just under 1,900 files and 400,000 lines, a fix of about 15 lines in a single function, one fail-to-pass test and 51 pass-to-pass tests. The means are pulled up by a few huge tasks, which is why the maximums matter too.

The math: what fraction of the code changes

In words: “the share of the codebase that the fix touches is the lines it edits divided by the lines there are.”

With the numbers: on average 32.8 / 438,000 = 0.000075, one line in about 13,000. For the median task, 15 / 400,000 = 0.0000375, one line in about 27,000.

In Python:

L_edit, L_code = 32.8, 438_000
rho = L_edit / L_code
f"{rho:.1e}", round(1 / rho, -3)  # → ('7.5e-05', 13000.0)
# the median task from Appendix A.5
round(400_000 / 15, -3)  # → 27000.0

Five more properties the paper lists

  • Real tasks: issues written by users, fixed by maintainers.
  • Continually updatable: the pipeline runs on any Python repository with minimal human work, so new tasks can come from issues created after a model's training data was collected.
  • Robust evaluation: every task has at least one fail-to-pass test, and 40% have at least two.
  • Edits across the codebase: no hint which function or file to change, unlike benchmarks that ask for one function or one blank to fill.
  • Many possible solutions: the reference fix is one answer; any patch that passes the tests counts.

Why it matters

A context of 438,000 lines does not fit in any model's context window, so the first problem is not writing code but choosing what to read, which is the job of fault localization. The coding agents lesson measures the same trade-off in miniature: dumping a repository costs every file's tokens, while searching first costs a few hundred.

2.4 SWE-bench Lite · original

Everyday picture

A full marathon tells you the most about a runner, but it is expensive to run every week. A 5 km time trial on the same course is a cheaper check.

Tiny example

SWE-bench Lite is 300 of the 2,294 tasks, 13% of them, chosen to be more self-contained and to focus on functional bug fixes. It covers 11 of the 12 repositories, in similar proportions. Counted in tasks, it is about an eighth of a full run.

Why it matters

SWE-agent, for one, reports its ablations on Lite and its headline results on the full set. One caution: the paper's own Appendix A.7, which should describe exactly how the 300 were selected, breaks off mid-sentence in every arXiv version (it reads “SWE-bench is” and stops), and the sentence in §2.4 pointing to the criteria is also cut short. The paper does not give the selection rule, so this page does not either.

3 SWE-Llama: fine-tuning CodeLlama for SWE-bench · original

Everyday picture

An apprentice who has read a lot of code, but never been handed a bug report and asked for a fix, tends to answer with something unrelated. A few thousand worked examples of “here is the report, here is the fix” teach the format.

Tiny example: the training data

The authors ran the same collection pipeline on 37 other Python repositories, disjoint from the 12 in the benchmark so no test task can leak into training, and dropped the requirement that a pull request add tests. That gave about 19,000 issue and pull request pairs. Excluding sequences longer than 30,000 tokens left about 10,000. Each training example is instructions, the issue, and the files the real fix edited; the target is the real patch (the paper calls it the “gold patch”).

CodeLlama-Python 7b, 13b SWE-bench-train37 other repos LoRA on the attentionprojections, r = 16 SWE-Llamawrites patches

Hover or tap a box.

SWE-Llama's recipe, drawn from §3 and Appendix B.1 of Jimenez et al. (2023).

Reading it: two inputs meet in the middle: a code model that can read long inputs, and pairs of issues and real fixes. The middle box trains only small added matrices inside the attention layers (LoRA, the LoRA companion takes it apart), chosen, the paper says, for memory efficiency. The 7b model trained in 20 hours on 4 A100 GPUs and the 13b in 47 hours on 8. The result writes patches directly.

Why it matters

At the time only CodeLlama among open models could read the long inputs needed, and off the shelf it produced placeholder or unrelated code. Fine-tuning made it competitive with Claude 2 in some settings, but it also taught it a habit that backfired, covered in §5: trained always to edit the files it was shown, it struggled when most files shown were irrelevant.

4 Experimental setup · original

A model cannot read 438,000 lines at once, so every baseline in the paper first chooses a few files, pastes them into one prompt, and asks for a patch in a single reply. No model runs code or looks around. This non-interactive setup is the baseline that later agents were measured against.

4.1 Choosing files: BM25 and the oracle · original

“We observe that in approximately 40% of instances, BM25 retrieves a superset of the oracle files for the 27,000-token context limit.”Jimenez et al. (2023), §4.1

Everyday picture

You can photocopy only 30 pages of the library for someone working on a problem. You pick pages by matching the words in their question against the index. If their question uses the same words as the page that holds the answer, you win; if not, you copy 30 pages of the wrong thing. A bigger photocopy budget helps you catch the right page, but it also buries it among more wrong ones.

Tiny example: two ways to choose files

BM25 (BM25, a classic keyword score from sparse retrieval) ranks every file by how well its words match the issue text, with the file's path prepended so a filename mentioned in the issue counts. The paper fills the prompt with top-ranked files up to a budget of 13,000, 27,000 or 50,000 tokens. The authors chose BM25 over dense retrieval because the queries and documents are very long and because matching natural-language issues to code is unusual territory for embedding models. Oracle retrieval simply hands the model the files the real fix edited. It is not realistic (an engineer doesn't know in advance which files need changing) and is used only for analysis.

Try it: eight files ranked by BM25 for one made-up issue, two of which the real fix edited. Raise the budget and watch which files make it into the prompt, how much of the oracle is found, and how much of the prompt is files the fix never touched.

Reading it: the files are listed in BM25's order, best match first, with their size in thousands of tokens (illustrative). Files are taken in order until the next one no longer fits. The two striped files are the ones the real fix edited. At 13k the prompt holds only a documentation page and a test file: recall 0. At 27k the first edited file arrives. At 50k both are in, but more than two thirds of the prompt is files the fix never touched. That is the trade the paper measures in Tables 2 and 3 below.

The math: three kinds of recall

In words: “Avg is the share of the edited files that were retrieved, averaged over tasks; All is the share of tasks where every edited file was retrieved; Any is the share where at least one was.”

With the numbers (illustrative): three tasks. Task 1 edits a.py and retrieval found it: 1, 1, 1. Task 2 edits two files and retrieval found one: 0.5, 0, 1. Task 3 edits one file and retrieval missed it: 0, 0, 0. So Avg = 0.5, All = 0.33, Any = 0.67. On the real benchmark at 27k tokens, the paper's Table 3 gives Avg 44.41%, All 39.83% and Any 51.27%: for almost half the tasks (48.7%) the prompt holds none of the right files.

In Python:

# (O_i, R_i) for three illustrative tasks: files the fix edited, files retrieved
tasks = [({"a.py"}, {"a.py", "b.py", "c.py"}), ({"x.py", "y.py"}, {"x.py", "z.py"}), ({"m.py"}, {"n.py", "p.py"})]
N = len(tasks)
avg = sum(len(R & O) / len(O) for O, R in tasks) / N
all_ = sum(O <= R for O, R in tasks) / N
any_ = sum(bool(R & O) for O, R in tasks) / N
round(avg, 2), round(all_, 2), round(any_, 2)  # → (0.5, 0.33, 0.67)
# SWE-bench at 27k tokens: tasks where BM25 found none of the edited files
round(100 - 51.27, 2)  # → 48.73

Hover the chart, or tab to it and use the arrow keys, to read BM25's recall at each budget.

Reading it: the x-axis is the budget for retrieved files, in thousands of tokens; the y-axis is recall in percent, from Table 3 of the paper. All three lines climb with the budget, as they must: more files, more chances to include the right ones. Even at 50k tokens, more than half the tasks (54.1%) are missing at least one file the fix edited.

Hover the chart, or tab to it and use the arrow keys, to read each model's resolved rate at each budget.

Reading it: the same x-axis; the y-axis is the percentage of tasks resolved, from Table 2 of the paper. Every line falls as the budget grows, the opposite of the recall chart above. More of the right files are in the prompt, and yet fewer tasks are resolved, because the right files arrive buried among more wrong ones. Both SWE-Llama models drop to 0 at 50k. (The paper's text refers to these resolved rates as Table 3; they are its Table 2.)

Why it matters

This pair of charts is one of the paper's most quoted lessons: a model's performance is limited not only by whether the answer is in its context but by how much else is there. The Lost in the Middle companion studies the same effect directly. For coding systems the conclusion was to stop choosing files once, up front, and let the model search and read as it goes, which is what agents do.

4.2 The prompt · original

Everyday picture

A request to a contractor who will visit only once: here is the complaint, here are the drawings of the rooms involved, here is an example of how to write up a change order, now write the change order.

Tiny example: the prompt's five parts

Instructions <issue> issue text </issue> <code>README.md, retrieved files,each between start and end markers <patch> an example patch </patch> “Respond with a single patch file”

Hover or tap each part of the prompt.

The prompt template of the paper's Appendix D.3, drawn as its parts. Based on Jimenez et al. (2023), Appendix D.3.

Reading it: read top to bottom, in the order the model reads. The pink part is nearly all of the length: the retrieved files. The example patch is there to show the format: the paper notes that models likely see few patch files in training, and many of their patches fail to apply. The request at the bottom asks for exactly one patch that can be applied with git apply. The model gets one reply, generated with greedy decoding, and no second chance.

Why it matters

The template shows what the baseline cannot do: ask a question, open another file, or run anything. Every piece of context must be chosen before the model starts. The authors' own list of future directions names the alternative, “agent-based approaches”, and a follow-up paper by several of the same authors (SWE-agent) built it.

4.3 The models · original

Everyday picture

If the drawings for a job don't fit through the door, the contractor never sees them. A model whose context window is too short for the relevant files starts with part of the task already lost.

Tiny example

Table 4 of Jimenez et al. (2023): each model's context window, and the share of tasks whose oracle files fit in it.
ModelMax tokensTasks where the oracle fits
ChatGPT-3.516,38558.1%
GPT-432,76884.1%
Claude 2100,00096.4%
SWE-Llama≥ 100,000≥ 94.8%

Reading it: the right column is a ceiling. ChatGPT-3.5 cannot even be shown the right files for 42% of tasks. Token counts are not directly comparable across models, the paper warns, because tokenizers differ: the same text is about 42% longer in Llama's tokens than in GPT-4's.

Why it matters

Context length was the gate for entering this benchmark at all in 2023. It still shapes how coding systems are built: whatever the window, a repository is bigger, so something must decide what to read.

5 Results · original

“Across the board, models struggle significantly to resolve issues.”Jimenez et al. (2023), §5

Everyday picture

A class sits a new, harder exam, and the best student scores 2%. That says more about the exam than about the students: it measures something the old exams did not.

Tiny example: the headline table

Reading it: each model has two bars, from Table 5 of the paper (BM25 retrieval, full benchmark). The solid bar is the percentage of tasks resolved, on an axis from 0 to 5%; the striped bar is the percentage of patches that even applied cleanly, on an axis from 0 to 100%. Look at the gap: between a quarter and a half of all patches apply, but only 0.17% to 3.79% of tasks are resolved. SWE-Llama's patches apply most often (it was trained on the format) yet resolve few. Claude 3 Opus, added in the 2024 revision, leads at 3.79%.

One small disagreement: the abstract and §5 give Claude 2 1.96%, while Table 5 gives 1.97%. Neither is off by much, but only one can be a count out of 2,294: 45 tasks is 1.96%, and 46 would be 2.01%.

# which whole number of tasks out of 2,294 gives each figure?
[round(100 * k / 2294, 2) for k in (44, 45, 46)]  # → [1.92, 1.96, 2.01]

Difficulty differs by repository

Table 19 of Jimenez et al. (2023), shaded: % resolved per repository with oracle retrieval. GPT-3.5 is ChatGPT-3.5, SL is SWE-Llama, and GPT-4 ran on a random 25% of tasks.

Reading it: the paper's Figure 4, redrawn from its Table 19 as a shaded table: one row per repository, one column per model, each cell the percentage resolved with oracle retrieval; darker means higher. The darkest row is psf/requests (up to 18.18%), whose codebase is among the smallest (119 files and 30,000 lines, from the paper's Table 11); seaborn is zero for everyone. The columns rise and fall together across rows, so difficulty is a property of the repository as much as of the model. GPT-4 ran on a random quarter of the tasks for budget reasons, so its column rests on few tasks per repository. The models also solve different tasks: Claude 2 solves 110 and SWE-Llama 13b 91, but only 42% of SWE-Llama's are among Claude 2's.

Images play a part too. Issues can embed screenshots, which a text model cannot see: 32% of matplotlib issues and 10% of seaborn issues contain an image, against 2% overall.

Difficulty rises with context; trimming it helps

Claude 2's resolved rate falls steadily as the total input gets longer (the paper's Figure 5), even with oracle files. To test whether the problem is the noise around the answer, the authors made an oracle-collapsed prompt: the oracle files, with everything more than 15 lines away from an edited line removed.

In words: “each edited region keeps its own lines plus a buffer of b lines above and b below; add that up over the regions.”

With the numbers (illustrative): a 1,000-line file with two edited regions of 4 and 6 lines keeps (4 + 30) + (6 + 30) = 70 lines, 7% of the file (ignoring regions that overlap or touch the file's ends). In the paper, collapsing raised GPT-4 from 1.3% to 3.4% and Claude 2 from 4.8% to 5.9% (Table 6), with Claude 3 Opus at 9.39%.

In Python:

b = 15
# e_h: lines edited in each region of one file (illustrative)
e = [4, 6]
L_kept = sum(e_h + 2 * b for e_h in e)
L_kept, L_kept / 1000  # → (70, 0.07)

Collapsing uses knowledge of the answer, so it is an analysis, not a method. What it shows is that the same model, given the same relevant code with less around it, does better: localization, not writing the fix, was a large part of the difficulty.

Four more findings

  • No sign of memorised fixes. Splitting tasks into those created before and after 2023 (Table 7), most models score about the same on both halves; only GPT-4 drops (1.96% to 0.0%), and it ran on a small subset. The authors read this as evidence that models are not simply recalling a later version of the code.
  • Fine-tuned models are brittle to what they are shown. SWE-Llama was trained with oracle files, which it learned always to edit. Given BM25 files, most of which should not be touched, it does poorly.
  • Patches beat whole files. Asked to rewrite each edited file in full instead of writing a patch, Claude 2 fell from 4.8% to 2.2% with oracle retrieval.
  • Model patches are small. Applied model patches change about half as many lines as the real fixes for the same tasks, and almost never more than one file (Table 8).
Selected rows of Table 8 of Jimenez et al. (2023): average size of applied patches with oracle retrieval, next to the real fixes for the same tasks.
PatchesTotal linesAddedRemovedFiles
Claude 219.64.21.91.0
Real fixes for the same tasks44.112.05.81.2
SWE-Llama 13b17.61.61.21.1
Real fixes for the same tasks37.810.04.41.1
All 2,294 real fixes74.522.310.51.7

Reading it: compare each model row with the row beneath it, which describes the real fixes for exactly the tasks where that model's patch applied. Claude 2's patches are 19.6 lines against 44.1, less than half. The last row, all real fixes, is bigger still: the tasks where a model's patch applied at all tend to have smaller real fixes. (“Total lines” is larger than added plus removed in every row; the paper does not say what else it counts.) The paper's own text compares 74.5 with 30.1, ChatGPT-3.5's figure, as if it were the models' average; the like-for-like comparison is the one in each pair of rows.

Why it matters

Every number here is a baseline that a model reading a pasted prompt could reach. Two findings pointed straight at what came next: context is the bottleneck, and models given a chance to check their work would do better. The authors say so in Appendix C.5, noting the value of letting a model run its fix against the tests before deciding to submit.

5.1 One task, up close · original

Everyday picture

A new employee told “the report prints the wrong heading when the setting is on” finds the function that prints the heading and makes it always print the right one for that setting. A senior engineer notices the function next door already reads the setting properly, and copies that pattern.

Tiny example: sphinx-doc__sphinx-8713

The issue: Sphinx's napoleon extension formats a docstring section called “Other Parameters” wrongly when the setting napoleon_use_param is True. With oracle retrieval, SWE-Llama 13b's input was 1,558 lines, or 20,882 tokens. It found the right function, _parse_other_parameters_section at line 684 of sphinx/ext/napoleon/docstring.py, and changed it to behave as if the setting were always True. The real fix checks the setting first, copying what the neighbouring _parse_parameters_section does. One of the pass-to-pass tests builds documentation with the setting False, and fails at once.

Why it matters

Across the 11 generations the authors studied (Appendix F has more), models wrote simple Python that ignored the codebase's existing helpers and style, and took a “greedy” path to the exact symptom in the issue. Real fixes often made structural changes that also prevented future bugs. The pass-to-pass tests caught this case; the paper's Discussion adds that passing tests alone doesn't guarantee a patch is as readable or efficient as a human's.

6 Related work and 7 Discussion · original

Everyday picture

A driving test on an empty car park checks each skill alone; a test in city traffic checks whether they come together.

What came before

Evaluation suites had either collected many small, separate tasks from different domains, or turned to interactive web tasks. Code benchmarks followed HumanEval: more languages, more tests, different edit sizes, library-specific problems. Software engineering research had studied program repair, bug localization and testing, but with far less code context. SWE-bench's claim is to combine these into one task, with minimal processing of real data, and to be easy to extend to other languages.

Limitations and what the authors hoped for

  • All tasks are Python.
  • The experiments are deliberately the simplest baselines; the authors encourage “agent-based approaches, tool augmented LMs”.
  • Passing tests is necessary but not sufficient: model code can be less comprehensive, efficient or readable than a human's. Appendix C.7 tries software engineering measures such as cyclomatic complexity: on one requests task, Claude 2's 6-line patch solved the problem but raised the complexity of a widely used class from 3 to 5, while the real 11-line fix touched a simpler function and added a new, specific error type.

The ethics statement notes that all 12 repositories have licences permitting this use (BSD, MIT, Apache, GPL 2.0 and custom ones, listed in Table 12), and that no human subjects or crowd workers were involved. Appendix E hopes the benchmark can also serve as a testbed for making AI-written code safe and faithful to what people intended.

Why it matters

The limitations list reads like a plan for the field's next two years: agents, more languages, and grading that looks beyond test results.

Appendix A: how a task is built and graded · original

“…a task is considered solved if all tests across FAIL_TO_PASS and PASS_TO_PASS pass.”Jimenez et al. (2023), Appendix A.4

Everyday picture

Before using an old exam question, a teacher checks the answer key really marks the model answer right and a blank page wrong. SWE-bench does the same check by machine, once per task, and records which tests are the ones that tell the difference.

Tiny example: the parts of a task, and the two test logs

Each task stores the problem statement P (the issue titles and bodies, plus only those comments written before the pull request's first commit, so no discussion of the solution leaks in), a pointer to the codebase C (the repository and the commit the pull request started from, kept in a mirror so history rewrites upstream can't break it), the gold patch δ (the non-test part of the pull request), and the test patch T (the part that touches files with “test” in their path). Comments on the pull request made before its first commit are saved as hints but not used in the paper's experiments.

To validate a task, the harness installs C in an environment built for that release of the repository, applies T, runs the tests (logpre), applies δ, runs them again (logpost), and sorts every test by how its status changed. For the leap-year pull request:

The two validation runs for the illustrative leap-year task, and the list each test lands in.
Testlogpre (old code)logpost (with δ)List
is_leap(1900) is FalsefailpassFAIL_TO_PASS
is_leap(2100) is FalsefailpassFAIL_TO_PASS
is_leap(2000) is TruepasspassPASS_TO_PASS
is_leap(2024) is TruepasspassPASS_TO_PASS
is_leap(2023) is FalsepasspassPASS_TO_PASS
days_in_february(2024) == 29passpassPASS_TO_PASS

Reading it: the list a test lands in is just its two statuses read left to right. The harness also records FAIL_TO_FAIL and PASS_TO_FAIL, but grading ignores them: a test that fails even with the real fix can't be required of a model. A task is discarded if the install or either run fails, if no test moves from fail to pass, or if logpre contains an ImportError or AttributeError, which the authors take as a sign that the tests depend on names only the fix introduces.

tests = ["1900", "2100", "2000", "2024", "2023", "feb2024"]
# log_pre and log_post: True means the test passed
pre = dict(zip(tests, [False, False, True, True, True, True]))
post = dict(zip(tests, [True, True, True, True, True, True]))
FAIL_TO_PASS = [t for t in tests if not pre[t] and post[t]]
PASS_TO_PASS = [t for t in tests if pre[t] and post[t]]
FAIL_TO_PASS, len(PASS_TO_PASS)  # → (['1900', '2100'], 4)

Diagram: grading one prediction

1–3: check out C, pickthe release's environment,install 4: apply test patch T 5–6: apply predicted patch;if it fails, repair once still fails:score 0 7: run the tests every FAIL_TO_PASS and PASS_TO_PASS testfound and passing: resolved

Hover or tap a step.

The evaluation procedure of Appendix A.4, redrawn from the paper's Figure 8 and its seven numbered steps. Based on Jimenez et al. (2023), Figure 8.

Reading it: follow the left column down. The first four steps never fail, because validation already proved them for this task. Step 5 is where a model's output can fail on format alone: if the patch's context lines don't match the file, the harness tries once to repair it (dropping unneeded context lines and recomputing the line counts in each @@ header) and gives 0 if it still won't apply. The box on the right is that exit. The final box is the resolved rule of §2.2: a test that is missing from the log counts as failed.

How often the repair was needed

The repair step mattered: Table 14 shows that for the closed models, 25% to 69% of the patches that applied needed it. ChatGPT-3.5 with BM25 produced 2,270 patches; 604 applied, and 363 of those only after repair (60.1%). SWE-Llama, trained on the format, needed it for 20% to 30%.

# Table 14: ChatGPT-3.5, BM25 13k: patches that applied, and how many needed repair
applies, fixed = 604, 363
round(100 * fixed / applies, 1)  # → 60.1

Why it matters

The whole pipeline makes grading a deterministic function of the patch: same patch, same verdict. That is what lets thousands of submissions be compared on a leaderboard. The BenchTask in the coding agents lesson stores the same two lists, and grade applies a patch to a fresh copy of the repository before running them, so nothing the agent did to its own workspace counts.

Appendix C: what an applied patch did · original

Everyday picture

“Failed” hides a lot. A student who got half the answers right and one who handed in a blank page both fail, but they need different help. The paper's Appendix C.5 splits every unresolved patch that applied into what it actually did.

Tiny example: six patches for one bug

Try it: pick a patch for the leap-year task. The grader runs the two fail-to-pass and four pass-to-pass tests and places the result in the paper's grid (its Table 22). The patches are deliberately wrong in different ways; find one for each of the six cells.


  

Reading it: the list shows each test's verdict for the chosen patch. The grid's columns count how many fail-to-pass tests pass (all, some, none) and its rows how many pass-to-pass tests pass; the outlined cell is where the patch lands. Only the top-left cell is resolved. Down the left column the issue is fixed but something else broke (“breaking resolved”), which is why the forgotten 400 rule loses. A patch that changes nothing lands in “no-op”; “always True” lands in “regression”: it fixes nothing and breaks 2023. The middle and bottom rows share names because the paper does not tell “some” from “none” of the pass-to-pass tests.

def outcome(f_pass, f_total, p_pass, p_total):
    # the grid of Table 22: columns by fail-to-pass, rows by pass-to-pass
    f = "all" if f_pass == f_total else ("none" if f_pass == 0 else "some")
    if p_pass == p_total:
        return {"all": "Resolved", "some": "Partially Resolved", "none": "No-Op"}[f]
    return {"all": "Breaking Resolved", "some": "Work in Progress", "none": "Regression"}[f]
# the patch that forgets the 400 rule: 2 of 2 fail-to-pass, 3 of 4 pass-to-pass
outcome(2, 2, 3, 4)  # → 'Breaking Resolved'
# a patch that changes nothing
outcome(0, 2, 4, 4)  # → 'No-Op'

What the models' patches did

Reading it: for the chosen model, each bar is the share of its applied patches (oracle retrieval) in one outcome, from Table 23 of the paper; the axis runs from 0 to 100%. Two bars dominate for every model: no-op and regression, patches that make not one fail-to-pass test pass. For Claude 2 they are 471 and 436 of 1,078 applied patches, 84% together. The three partial outcomes are small. The authors read this as models often understanding the issue but lacking a view of how the changed code is used elsewhere.

Two places where this appendix disagrees with the rest of the paper. The “applied” counts here (1,078 for Claude 2, 1,196 for SWE-Llama 13b) are smaller than the counts in Table 14 for the same setting (1,441 and 1,532), which match the apply rates of Table 18; the paper does not explain the difference. And the text says no-ops are “60% to 70%” of the patches that pass no fail-to-pass test; that holds for ChatGPT-3.5 and both SWE-Llama models, but for Claude 2 it is 471 of 907 (52%) and for GPT-4 30 of 59 (51%).

# Table 23, Claude 2: resolved, breaking resolved, partially, work in progress, no-op, regression
counts = [110, 26, 15, 20, 471, 436]
sum(counts)  # → 1078
# no-op share among patches that pass no fail-to-pass test
round(471 / (471 + 436), 2)  # → 0.52

Why it matters

The grid is a sharper instrument than one pass rate. A system stuck in “no-op” is not finding the bug; one stuck in “breaking resolved” finds it but doesn't know the code around it. Those failures call for different fixes: better search in the first case, running the existing tests in the second. The authors end the appendix by pointing at exactly that: an environment where a model can run its fix against the tests before submitting.

What happened next

In the paperWhat came laterLearn it
One prompt, one reply, files chosen by BM25Agents that search, open files, edit and run tests in the repository: SWE-agent (Yang et al., 2024, arXiv:2405.15793) resolved 12.47% of the full benchmark, against 3.79% for the best retrieval baselineSWE-agent companion
300-task Lite subsetUsed for cheaper comparisons, such as SWE-agent's ablationsSWE-agent's ablations
Fail-to-pass and pass-to-pass tests, resolved rateA common way to grade agent-written code changesa three-task version
Collection that keeps running on new issuesA defence against contamination, as the original tasks become publicbenchmarks lesson
Localization as the bottleneckSearch tools built for models, and reading code in small windowsWorkspace.search

Glossary

Every term with hover guidance on this page, in one place.