Over the last 90 days, our one-engineer team merged 1,456 pull requests and deployed to production 155 times. That didn't happen by abandoning the software development lifecycle. It happened by compressing every phase of it into an AI loop with a human gate at the exit — and letting DORA metrics referee the whole machine. Here's the working system, with the real numbers.
Source: the aictrl-dev/aictrl repository and our own review telemetry, 2026-04-19 → 2026-07-18. Methodology and caveats are stated inline throughout — this is an n=1 case study, not a benchmark.
In February 2025, Andrej Karpathy named the mood: "vibe coding" — fully give in to the vibes, embrace exponentials, and forget that the code even exists. In November 2025, Collins Dictionary made it a Word of the Year. By early 2026 the backlash essays were everywhere, and even Karpathy had moved on — as CodeRabbit's history of the term records, he refined the thesis into what he and others call "agentic engineering": "programming via LLM agents is increasingly becoming a default workflow for professionals, except with more oversight and scrutiny." The mega-prompt gave way to strategic decomposition.
Meanwhile the process world held a funeral of its own. Boris Tane declared the SDLC dead: "The entire lifecycle, the one we've built careers around, the one that spawned a multi-billion dollar tooling industry, is collapsing in on itself." And at the other extreme, plenty of engineering organizations concluded that nothing fundamental changed — keep the sprint rituals, keep the review checklist, add a Copilot licence.
Our experience running an AI-heavy delivery machine for the past year is that both funerals are premature. The classic phases — requirements, design, code, review, release — all survived. What changed is their shape. Each phase collapsed into a loop: AI iterating against AI at machine speed, with a deterministic backstop where possible. And human involvement, instead of being spread thin across every phase, concentrated into two gates: what to build, and what to ship.
"AI doesn't fix a team; it amplifies what's already there. Strong teams use AI to become even better and more efficient. Struggling teams will find that AI only highlights and intensifies their existing problems."
DORA's framing matters here: the loops and gates are the "what's already there." AI didn't build our process. It made the process the only thing standing between us and 1,456 unreviewed merges.
Nothing about writing specs is new. What's new is that the spec conversation has a tireless counterparty. An idea goes in; an AI plays analyst and critic against it — is it right-sized, is it testable, does it align with the roadmap, what's out of scope? The loop runs until the spec earns one of three verdicts: ready, refine, or split. What comes out the other side is not a chat transcript: it's backlog issues with acceptance criteria, an ADR or RFC in the team wiki, and design notes. Durable artifacts, in version control or the wiki, exactly where a 2015-era process would have put them.
The human gate at the exit is spec approval — and it's the cheapest, highest-leverage decision in the whole pipeline. The industry is converging on the same conclusion from several directions. GitHub built an entire open-source toolkit — Spec Kit — around the observation that ad-hoc prompting produces code that "looks right, but doesn't quite work," and that the fix is making a versioned spec, not the prompt, the unit of work. Harper Reed's much-shared LLM codegen workflow is the same shape discovered independently: "Brainstorm spec, then plan a plan, then execute using LLM codegen. Discrete loops."
One honest meta-note: this article went through the same loop. Its design spec was drafted, critiqued, and approved in our wiki before a word of prose existed. The spec loop cost about thirty minutes. It's the phase most teams skip, and it's the one that decides whether the next phase produces convergent iterations or expensive noise.
The 2025 Stack Overflow survey (49,000+ developers) found the single biggest AI frustration, cited by 66%, is "AI solutions that are almost right, but not quite" — and 45% say debugging AI-generated code is more time-consuming. That "almost right" tax is real, and you only have two ways to pay it: with human attention, which recreates the old bottleneck, or with more machines.
We pay it with machines. An agent writes the code. An AI reviewer — with full repository context — reads the diff and files structured findings. The agent fixes what's real, pushes, and the review runs again on the new commit. The loop repeats until the reviewer runs clean, with CI and deterministic static checks as the backstop underneath. The human is not in this loop.
code_review_runs), aictrl-dev/aictrl repository.
The numbers behind the loop, from our own telemetry over the same 90 days: 2,214 review runs across 988 PRs; a median of 3 findings per run (mean 4.4, p90 = 10). Roughly 39% of findings had their cited code changed before merge (n=9,171 findings with actionability computed) — and in the subset a human or agent explicitly triaged (n=1,653), 79% were judged true positives. Not every finding is gold. Enough are that two machine-speed rounds catch what would otherwise wait days for a human reviewer's attention.
Here's the contrarian part. That same Stack Overflow survey found only 10.2% of developers mostly use AI for committing and reviewing code — a further 22.6% use it partially, and 58.7% have no plans to use it for this task at all. The industry adopted AI where it feels good (generation) and has been slowest to adopt it where the leverage is (verification). DORA's researchers saw this coming — their interviews found time saved generating code gets re-allocated to verification overhead: "Reviewing [another's] code is so much harder than writing it. AI tools are increasing the rate at which people can churn out code that needs to be reviewed." If you accelerate generation without accelerating verification, you haven't compressed the SDLC. You've moved its queue.
The trust numbers explain why this gate survives everything else being automated: 84% of developers use or plan to use AI tools, but trust in the accuracy of AI output fell from 40% to just 29% year over year. Adoption went up while trust went down. The gap between those two curves is exactly where a human gate belongs.
But holding the gate can't mean reading every line. At 1,456 PRs a quarter, line-by-line human review isn't diligence, it's fiction — the kind that gets rubber-stamped at 5pm. So the gate works on prepared evidence: by the time a PR reaches the human, it carries an AI-written summary of what changed and why, a risk framing, and the full review trail — what the AI reviewer found in round one, what got fixed by round two, what was dismissed and on what grounds. The human reviews the decision, not the diff. Does this change belong in the product? Minutes, not hours.
That distribution is the human gate's report card: it only holds at this speed because the machine did the reading first.
Where our gate slipped — and how we noticed
Full disclosure: while pulling the data for this article we found a contiguous block of ~20 PRs (a mid-July design-system series) that merged with no AI review at all — confirmed absent in both GitHub and our telemetry. Overall review coverage in the window is roughly 68% of merged PRs, not 100%. That's the real lesson about gates: they don't fail loudly, they fail silently, and the only reason we know this one slipped is that the scoreboard is instrumented. A gate nobody measures quietly stops existing.
After the drama of three loops, deployment should be an anticlimax, and it is: merge to main triggers the pipeline, and the pipeline shipped to production 155 times in 90 days — about 12 deploys a week, every one a green run of the same workflow. (The sandbox environment deploys far more often still — north of 85 times a week — because every loop iteration wants a place to prove itself.) There is no release committee and no deploy window. All the heroics moved upstream into the loops; by the time code reaches this phase, shipping it is the least interesting thing that happens to it.
DORA's current five-metric definition pairs this throughput with stability — change fail rate, recovery time, rework rate. We're publishing our throughput keys in this piece and not our stability keys, for an unglamorous reason: our change-failure and recovery numbers aren't yet instrumented to a standard we'd defend in public. Saying so out loud is part of the same discipline the rest of this article argues for.
In 2024, DORA delivered the industry's least welcome finding: "the effect on delivery throughput is small, but likely negative (an estimated 1.5% reduction for every 25% increase in AI adoption). The negative impact on delivery stability is larger (an estimated 7.2% reduction…)" (2024 report, p. 40). A year later, the 2025 report saw throughput flip positive while stability stayed negative, and named the mechanism: "AI accelerates software development, but that acceleration can expose weaknesses downstream. Without robust control systems… an increase in change volume leads to instability."
"Robust control systems" is the load-bearing phrase. The loops in this article are those control systems. Field data shows what their absence looks like: Faros's study of 10,000+ developers found AI-assisted teams shipping 47% more PRs per day — with 154% larger PRs, 91% longer code reviews, and 9% more bugs per developer. Volume went up; the bottleneck moved to the un-automated phase. That's not compression, that's displacement.
There's a subtler cost to the "SDLC is dead" framing, and it's the strongest argument for keeping named phases: a blob has no seams to measure. Collapse development into one undifferentiated agent loop — intent in, deploy out — and you lose the instrument-attachment points, and with them the ability to improve anything on purpose. A defined workflow keeps every classic flow question answerable: how long does an approved spec wait before an agent picks it up? How long is the queue at the sign-off gate? How much rework happens before release versus after? What does each phase cost — once measured in engineer-hours, now increasingly in tokens and machine-minutes? These are twenty years of questions we learned to ask about human teams. AI didn't retire a single one of them; it changed the unit prices.
And because the pipeline is instrumented end to end, we can publish the unit prices — with one scoping note that is itself the point. The delivery and review figures above are server-side and complete for the full quarter, but token cost comes from on-disk agent transcripts, and those get pruned. Only June 2026 is a fully-covered calendar month in the local history — all 30 days present — so we price the loops on June rather than defend a 90-day total whose early weeks are half-gone. (A cost metric you can't fully reconstruct is exactly the kind of gap a named, instrumented phase is supposed to surface.) In June, the local agents running the spec and coding loops consumed 12.05 billion tokens, which prices out at $9,009 at Anthropic API list rates — a blended ~$0.75 per million tokens across the model mix actually used, measured with the open-source ccusage across every Claude Code transcript on the operator's machine. (The blend is that low because of the cache-read share below; pricing the same volume at a single model's input rate would give a very different — and wrong — number.) Against June's 583 merged PRs, that's about $15 per merged PR for the local loops. The remote review loop adds roughly $1.6 per review round at z.ai list prices (GLM-5.1: $1.40/M input, $0.26/M cached, $4.40/M output; ~2.7 million tokens per round) — so an all-in merged PR lands around $18 at the two providers' list prices.
Two honesty notes on those figures. First, "at list prices" is a counterfactual: both loops actually run on flat-rate subscriptions, so the marginal cash cost of one more review round is close to zero — which is exactly why rework moved into the machine. Second, 96% of the local tokens are cache reads: the loops are economically viable because prompt caching makes re-reading the same repository on every round nearly free. Token cost is now a per-phase, per-experiment metric like any other on the scoreboard — which is the whole point of keeping the phases named.
Figure 1 back in Part 3 is exactly this kind of instrument: pre-release rework made visible and priced. A review round averages about 18 machine-minutes in our telemetry — rework at its cheapest, caught before merge. DORA's change-fail and rework rates measure the same rework after release, where the unit price is incidents and human attention. "Shift quality left" only works — and only provably works — if the workflow has a left to shift toward.
Phase boundaries are also what make experiments possible: change one stage, hold the others constant, read the delta off the scoreboard. Before trusting the review loop we ran it through a head-to-head evaluation on 212 of our own PRs; it earned its place in the pipeline with data, and it can lose its place the same way. A blob can't be A/B tested. A workflow can.
And the necessary honesty about what this case study is: it's an n of 1. One product, one human engineer (98.7% of PRs merge under one login, the rest under a second contributor) plus a fleet of AI agents that commit under that same account — which also means GitHub metadata cannot cleanly separate "AI-authored" from "hand-authored" work, and we won't pretend it can. We claim existence, not superiority: a solo-operator SDLC sustaining ~112 merged PRs and ~12 production deploys a week for a quarter, with the loops and gates described above, is possible and measurable. Your numbers will differ. The shape, we suspect, won't.
Skip the spec loop and the coding loop iterates on a moving target: every ambiguous requirement becomes another regenerate-from-scratch round. GitClear's January 2026 maintainability research — over 600 million commits — found within-commit copy/paste up 41%, code-block duplication up 81% and refactoring line moves down 70%, with two-week code churn up 15%. Speed without convergence.
Skip the AI review loop and the 66% "almost right" tax lands on humans — Faros's field numbers above (+91% review time) are what that looks like at scale. The human gate becomes the queue DORA warned about.
Skip the human gate and you're shipping code that nearly half of developers actively distrust, unread. That experiment ran industry-wide in 2025 under the name vibe coding; the retrospectives are not kind — and we've written before about where velocity without outcomes leads.
Skip the scoreboard and you won't know which of the above already happened. Our own unreviewed-PR block is the proof: the gate slipped silently, and only the metrics noticed. DORA's team frames AI success as cultivating conditions, not buying tools — working in small batches, strong version control practices, quality internal platforms. The scoreboard is how you find out whether your conditions are actually cultivated.
The SDLC didn't die. Its phases compressed into loops, its humans concentrated into gates, and its metrics stopped being a quarterly slide and became the control loop that keeps a machine-speed pipeline honest. The engineering craft didn't disappear either — it moved up a level: designing the loops, holding the gates, watching the panel, and improving the machine deliberately, one measured experiment at a time.
The concrete next step, if you want to test the thesis on your own team this week: measure three numbers you probably don't have — how many review rounds a PR takes before merge, your PR-open-to-merge lead-time distribution, and how long finished work sits queued at your human gates. If you can't produce them, that's your Step 0. If you can, you now have a baseline — and the first phase worth compressing is whichever one your baseline says is the queue.
How aictrl.dev helps
Disclosure: aictrl.dev builds the tooling we used to run and measure the pipeline in this article — AI code review with structured findings and verdict tracking, workflow orchestration for the loops, and delivery analytics for the scoreboard — so we have a commercial interest in the practices described here. The cited research is independent, and every practice above runs on any stack: the loops are a process shape, not a product feature. If you want the instrumented version rather than building it yourself, see how aictrl.dev can help.
First-party data: merged-PR, lead-time and deploy figures from the aictrl-dev/aictrl repository via GitHub CLI (window [2026-04-19T00:00Z, 2026-07-18T00:00Z) — start inclusive, end exclusive; date-chunked queries, totals cross-checked. GitHub's search API, whose merged: range is inclusive at both ends, returns 1,463 for the same window); review-loop and acceptance figures from aictrl review telemetry in BigQuery (code_review_runs, code_review_findings, deduplicated to latest row per finding). Token and cost figures are scoped to June 2026 — the only fully-covered calendar month (30/30 days) in the local ccusage history, since on-disk agent transcripts are pruned and the April–May local record is incomplete; delivery and review figures elsewhere remain the full 90-day window. June local totals from ccusage at Anthropic API list prices (a counterfactual — actual billing is subscription-based), divided by June's 583 merged PRs (the ~$15/PR figure divides the full local bill, of which a small fraction — ~9% over the 90-day sample — is non-product work). Remote review-run tokens from code_review_runs.token_usage, recorded on ~19% of June runs (829 runs, so the per-round mean is indicative and the review-loop total is a floor), priced at official z.ai GLM-5.1 list rates ($1.40/$0.26/$4.40 per M input/cached/output); runs without model metadata are priced at GLM-5.1 rates, the review executor's model family. Known limitations: n=1 team; AI-vs-human PR authorship not separable (agents commit under the operator's account); review coverage ≈68% of merged PRs (approximate); lead time measured PR-open → merge; stability keys (change fail rate, recovery) not published; token figures exclude cloud executor tasks other than code review and CI compute.
Published 2026-07-18 · Analysis by Bulat at aictrl.dev, co-authored with AI