AI Coding Research / February 2026
Frontier AI models that ace coding benchmarks fail dramatically at real software evolution tasks. GPT-5 drops from 65% to 21% when moving from isolated fixes to long-horizon multi-file changes.
For two years the industry has compressed "can AI write code?" into a single number: the share of SWE-Bench issues an agent resolves. That number has climbed relentlessly — past 50%, past 70%, into leaderboard bragging rights. But SWE-Bench, like nearly every coding benchmark before it, tests a narrow slice of the job: take one isolated bug report, change a handful of lines, make a failing test pass. Production engineering looks almost nothing like that.
SWE-EVO, introduced by Tue Le and colleagues in late 2025, is the first benchmark built to measure the work that actually consumes teams: long-horizon software evolution. Instead of single-issue patches, each task asks an agent to deliver an entire release's worth of change — interpreting high-level requirements, coordinating edits across many files, and preserving existing behaviour across multiple iterations.
The 48 tasks were mined from the release notes of seven mature open-source Python projects — among them scikit-learn, pydantic, dvc and dask. For each one the authors took a tagged release, defined the problem statement as the release-note delta to the next version, and kept only instances with at least one failing test that the change must turn green. The scale is what separates a benchmark from a job: an average of 21 files touched per task, validated against suites averaging 874 tests per instance.
Tasks
48
Real evolution scenarios from 7 mature Python projects
Avg Files
21
Files modified per task — true multi-file reasoning
Tests
874
Per instance validating functionality preservation
SWE-EVO also reports two numbers rather than one. Resolved Rate is strict and binary: an agent scores only when every target test passes and nothing previously working breaks. Fix Rate awards partial credit for the fraction of target tests turned green — but still zeroes any run that regresses existing behaviour. On long-horizon work, "almost done" and "broke the build" are very different outcomes, and a single pass/fail score hides both.
Source: SWE-EVO Paper (v1)
Here is the headline. GPT-5 — the best performer in the comparison below — resolves 65% of SWE-Bench Verified issues. Pointed at SWE-EVO, the same model resolves just 20.8%. Not a few points of degradation: a three-fold collapse. The capability that reads as "senior engineer" on isolated tickets reads as "confused intern" the moment the task spans a real release.
The gap is wide because the two benchmarks reward different skills. SWE-Bench Verified rewards localised pattern-matching — find the bug, mirror the surrounding code, pass the test. SWE-EVO rewards sustained reasoning over a large, interdependent change: holding 21 files in working memory, sequencing edits so nothing breaks midway, and not losing the thread across dozens of tool calls. Today's models are extraordinary at the first and unreliable at the second.
SWE-Bench Verified (isolated) vs SWE-EVO (long-horizon)
Same model. Same tasks. Different scope. The gap between isolated fixes and sustained multi-file evolution is enormous.
Source: SWE-EVO Paper, Vals.ai SWE-Bench
The collapse is not a GPT-5 quirk — it is universal. Every one of the eight models compared here loses between 25 and 52 points moving from isolated fixes to evolution. O3 falls from 58.4% to 6.25%. DeepSeek-R1 from 57.6% to 8.33%. Even the narrowest gap, Kimi-K2's, is a 25-point drop. The long-horizon ceiling is low for everyone, and a model's SWE-Bench rank barely predicts where it lands.
The paper's most useful finding is that models don't merely fail more on SWE-EVO — they fail differently. When the authors categorised GPT-5's failures, roughly 60% were instruction-following errors: it misread long, nuanced release notes rather than fumbling the edit or the tooling. Smaller models such as GPT-5-nano failed the opposite way, racking up tool-use and syntax errors — losing control of the interface before semantics ever came into play. Open-weight Kimi-K2 showed the reverse signature again: around 70% "incorrect implementation" with almost no tool-use trouble (solid interface control, weaker reasoning), while reasoning-tuned DeepSeek-R1 tended to get stuck in loops and exit early.
That split has a direct operational consequence. With a frontier model, the bottleneck is the specification — hand it a vague brief and it will confidently build the wrong thing. With a cheaper or smaller model, the bottleneck is the scaffolding — it needs tighter tools, smaller steps and more guardrails just to stay on the rails.
The harness — the framework that turns a model's tokens into shell commands, file edits and test runs — is a real variable, though an uneven one. Run GPT-4.1 under SWE-Agent's CLI-style access instead of OpenHands' containerised CodeActAgent and its score quintuples, from 2.08% to 10.42%. Yet GLM-4p5 and Qwen3-Coder are a wash across both, and DeepSeek-R1 actually does worse under SWE-Agent.
The paper itself stops short of crowning a winner — both harnesses were capped at 100 iterations, and the authors attribute most of the spread to the models rather than the framework. But the lesson for teams holds: when a model is near the edge of its competence, the execution environment around it can swing results several-fold. Below the frontier, your tooling architecture isn't a detail — it's part of the model's effective capability.
Biggest Swing
5x
GPT-4.1: 2.08% → 10.42% with SWE-Agent
Exception
-20%
Deepseek-R1 performs worse on SWE-Agent
What's a harness? The agent framework executing LLM actions. SWE-Agent uses CLI-based shell access; OpenHands runs in isolated containers. Your harness architecture matters as much as model selection.
Source: SWE-EVO Paper (Table 2)
Important: Claude models were not included in SWE-EVO. Here's performance on related benchmarks:
aictrl.dev analysis — based on GPT-5's 3x drop, Claude Opus 4.5's estimated SWE-EVO is ~25-30% (not independently evaluated).
A model scoring 65% on SWE-Bench Verified may only achieve 21% on real software evolution. The gap between isolated fixes and sustained multi-file work is enormous. Treat a headline leaderboard score as a ceiling on demos, not a forecast of production throughput.
AI excels at bounded, single-file changes. A 2-day refactor should become 15-20 well-scoped AI tasks, not one "please refactor this system" prompt. Decomposition has stopped being the busywork you do before delegating — it is now the work that decides whether the agent succeeds at all.
The same model can swing from 2% to 10% based on execution framework. Your tooling architecture is as important as model selection. Evaluate the agent framework, sandbox and tool surface as deliberately as you evaluate the model behind them.
Frontier models fail on instruction following (misinterpreting nuanced specs). Weaker models fail on tool use and syntax. Train your team accordingly — match the review effort to the failure mode: spec-check the frontier models, sandbox-and-step the smaller ones.
The 21% vs 65% gap is fundamentally about context. Engineers who excel with AI tools have learned to manage context boundaries explicitly. The scarce skill is no longer typing code — it's deciding what the model sees, in what order, and when to checkpoint.
The takeaway from SWE-EVO is not "AI can't code." It plainly can — inside boundaries. The benchmark's real message is that the line between a demo and a production system is long-horizon context, and that drawing that line well — decomposing work, managing context, verifying behaviour across many files — is becoming the core of the engineering job rather than an afterthought to it.