AI Coding Research / February 2026

SWE-EVO Benchmark:
The 3x Performance Collapse

Frontier AI models that ace coding benchmarks fail dramatically at real software evolution tasks. GPT-5 drops from 65% to 21% when moving from isolated fixes to long-horizon multi-file changes.

Performance drop 3x from 65% to 21%
1

What is SWE-EVO?

For two years the industry has compressed "can AI write code?" into a single number: the share of SWE-Bench issues an agent resolves. That number has climbed relentlessly — past 50%, past 70%, into leaderboard bragging rights. But SWE-Bench, like nearly every coding benchmark before it, tests a narrow slice of the job: take one isolated bug report, change a handful of lines, make a failing test pass. Production engineering looks almost nothing like that.

SWE-EVO, introduced by Tue Le and colleagues in late 2025, is the first benchmark built to measure the work that actually consumes teams: long-horizon software evolution. Instead of single-issue patches, each task asks an agent to deliver an entire release's worth of change — interpreting high-level requirements, coordinating edits across many files, and preserving existing behaviour across multiple iterations.

The 48 tasks were mined from the release notes of seven mature open-source Python projects — among them scikit-learn, pydantic, dvc and dask. For each one the authors took a tagged release, defined the problem statement as the release-note delta to the next version, and kept only instances with at least one failing test that the change must turn green. The scale is what separates a benchmark from a job: an average of 21 files touched per task, validated against suites averaging 874 tests per instance.

Tasks

48

Real evolution scenarios from 7 mature Python projects

Avg Files

21

Files modified per task — true multi-file reasoning

Tests

874

Per instance validating functionality preservation

SWE-EVO also reports two numbers rather than one. Resolved Rate is strict and binary: an agent scores only when every target test passes and nothing previously working breaks. Fix Rate awards partial credit for the fraction of target tests turned green — but still zeroes any run that regresses existing behaviour. On long-horizon work, "almost done" and "broke the build" are very different outcomes, and a single pass/fail score hides both.

Source: SWE-EVO Paper (v1)

2

The Performance Gap

Here is the headline. GPT-5 — the best performer in the comparison below — resolves 65% of SWE-Bench Verified issues. Pointed at SWE-EVO, the same model resolves just 20.8%. Not a few points of degradation: a three-fold collapse. The capability that reads as "senior engineer" on isolated tickets reads as "confused intern" the moment the task spans a real release.

The gap is wide because the two benchmarks reward different skills. SWE-Bench Verified rewards localised pattern-matching — find the bug, mirror the surrounding code, pass the test. SWE-EVO rewards sustained reasoning over a large, interdependent change: holding 21 files in working memory, sequencing edits so nothing breaks midway, and not losing the thread across dozens of tool calls. Today's models are extraordinary at the first and unreliable at the second.

GPT-5 Performance Comparison

SWE-Bench Verified (isolated) vs SWE-EVO (long-horizon)

SWE-Bench Verified
65.0%
SWE-EVO
20.8%
0% 25% 50% 75% 100%
!

Same model. Same tasks. Different scope. The gap between isolated fixes and sustained multi-file evolution is enormous.

Source: SWE-EVO Paper, Vals.ai SWE-Bench

3

Full Model Comparison

The collapse is not a GPT-5 quirk — it is universal. Every one of the eight models compared here loses between 25 and 52 points moving from isolated fixes to evolution. O3 falls from 58.4% to 6.25%. DeepSeek-R1 from 57.6% to 8.33%. Even the narrowest gap, Kimi-K2's, is a 25-point drop. The long-horizon ceiling is low for everyone, and a model's SWE-Bench rank barely predicts where it lands.

SWE-EVO (long-horizon)
SWE-Bench Verified (isolated)
GPT-5
-44%
GPT-5-mini
-49%
O3
-52%
Deepseek-R1
-49%
Qwen3-Coder
-41%
GLM-4p5
-38%
Kimi-K2
-25%
GPT-4.1
-29%
0% 20% 40% 60% 80%

Source: SWE-EVO Paper (Table 1), Vals.ai SWE-Bench Verified

The paper's most useful finding is that models don't merely fail more on SWE-EVO — they fail differently. When the authors categorised GPT-5's failures, roughly 60% were instruction-following errors: it misread long, nuanced release notes rather than fumbling the edit or the tooling. Smaller models such as GPT-5-nano failed the opposite way, racking up tool-use and syntax errors — losing control of the interface before semantics ever came into play. Open-weight Kimi-K2 showed the reverse signature again: around 70% "incorrect implementation" with almost no tool-use trouble (solid interface control, weaker reasoning), while reasoning-tuned DeepSeek-R1 tended to get stuck in loops and exit early.

That split has a direct operational consequence. With a frontier model, the bottleneck is the specification — hand it a vague brief and it will confidently build the wrong thing. With a cheaper or smaller model, the bottleneck is the scaffolding — it needs tighter tools, smaller steps and more guardrails just to stay on the rails.

4

Harness Choice: SWE-Agent vs OpenHands

The harness — the framework that turns a model's tokens into shell commands, file edits and test runs — is a real variable, though an uneven one. Run GPT-4.1 under SWE-Agent's CLI-style access instead of OpenHands' containerised CodeActAgent and its score quintuples, from 2.08% to 10.42%. Yet GLM-4p5 and Qwen3-Coder are a wash across both, and DeepSeek-R1 actually does worse under SWE-Agent.

The paper itself stops short of crowning a winner — both harnesses were capped at 100 iterations, and the authors attribute most of the spread to the models rather than the framework. But the lesson for teams holds: when a model is near the edge of its competence, the execution environment around it can swing results several-fold. Below the frontier, your tooling architecture isn't a detail — it's part of the model's effective capability.

OpenHands (container-based)
SWE-Agent (CLI-based)
GPT-4.1
+5x
GPT-5
+11%
Kimi-K2
+12%
GLM-4p5
Tie
Qwen3-Coder
Tie
Deepseek-R1
-20%
O3
+50%
0% 5% 10% 15% 20%+

Biggest Swing

5x

GPT-4.1: 2.08% → 10.42% with SWE-Agent

Exception

-20%

Deepseek-R1 performs worse on SWE-Agent

!

What's a harness? The agent framework executing LLM actions. SWE-Agent uses CLI-based shell access; OpenHands runs in isolated containers. Your harness architecture matters as much as model selection.

Source: SWE-EVO Paper (Table 2)

Claude / Anthropic Models

Important: Claude models were not included in SWE-EVO. Here's performance on related benchmarks:

80.9%
SWE-Bench Verified
Claude Opus 4.5 — #1 (Feb 2026)
74.6%
SWE-Bench Std
Vals.ai harness
54%
Claude Code
SWE-Bench-Pro 30d
???
SWE-EVO
Not evaluated

aictrl.dev analysis — based on GPT-5's 3x drop, Claude Opus 4.5's estimated SWE-EVO is ~25-30% (not independently evaluated).

Source: Vals.ai SWE-Bench, MarginLab Claude Code Tracker

5

Key Takeaways for Engineering Leaders

1

Benchmarks Lie (By Omission)

A model scoring 65% on SWE-Bench Verified may only achieve 21% on real software evolution. The gap between isolated fixes and sustained multi-file work is enormous. Treat a headline leaderboard score as a ceiling on demos, not a forecast of production throughput.

2

Task Decomposition is Non-Negotiable

AI excels at bounded, single-file changes. A 2-day refactor should become 15-20 well-scoped AI tasks, not one "please refactor this system" prompt. Decomposition has stopped being the busywork you do before delegating — it is now the work that decides whether the agent succeeds at all.

3

Harness Choice Matters More Than You Think

The same model can swing from 2% to 10% based on execution framework. Your tooling architecture is as important as model selection. Evaluate the agent framework, sandbox and tool surface as deliberately as you evaluate the model behind them.

4

Strong Models Fail Differently

Frontier models fail on instruction following (misinterpreting nuanced specs). Weaker models fail on tool use and syntax. Train your team accordingly — match the review effort to the failure mode: spec-check the frontier models, sandbox-and-step the smaller ones.

5

The "10x Engineer" Becomes "10x Context Manager"

The 21% vs 65% gap is fundamentally about context. Engineers who excel with AI tools have learned to manage context boundaries explicitly. The scarce skill is no longer typing code — it's deciding what the model sees, in what order, and when to checkpoint.

The takeaway from SWE-EVO is not "AI can't code." It plainly can — inside boundaries. The benchmark's real message is that the line between a demo and a production system is long-horizon context, and that drawing that line well — decomposing work, managing context, verifying behaviour across many files — is becoming the core of the engineering job rather than an afterthought to it.

Sources

SWE-EVO: Benchmarking Long-Horizon Coding — arXiv, December 2025 (v1) SWE-Bench Leaderboard — Vals.ai Claude Code Tracker — MarginLab