There are two ways to put AI agents to work, and they are not two settings of one dial. One is an engineer steering a harness turn by turn. The other is a workflow that pulls its own tasks, picks its own models and skills, and ships artifacts unattended. They are different machines — different economics, different failure modes, different management models. This is the case for moving everything that can move into the second machine, the evidence for where that stops, and the rule for telling the two apart.
For engineering leaders — the one-screen version
| The claim | Attribution requires boundaries. A chat session has no primary key, so nothing measured inside one can be defended, adopted, or rolled back by anyone but the person who ran it. |
| The test | Do two or three people already do this task? Plurality of practitioners proves a spec exists, even if nobody wrote it down. |
| The constraint | Not model choice, not prompting. Verification. Autonomy is bounded by what you can check without a human looking. |
| First 30 days | Pick one recurring task type. Write down its acceptance criteria — especially the negative rules. Retain inputs and outputs from today, whether or not you automate. |
| The outcome | Task types with a measured success rate you can move on purpose — instead of personal productivity that evaporates when the person does. |
In an interactive harness, the engineer silently supplies four things at every turn: context, taste, verification, and accountability. Nobody writes them down because nobody has to — they are in the room. Every step toward autonomy is the same move repeated: take one of those four and externalise it into an artifact that survives the session.
That is why the automated workflow is not a scaled-up version of the interactive one. It is the same work with the human's silent contributions made explicit — and most of the difficulty is in the making-explicit, not in the automating.
| Interactive harness loop | Automated workflow | |
|---|---|---|
| Who initiates | An engineer, ad hoc | The system pulls from a queue |
| Unit of record | A chat session | A task with an ID |
| Steering | Human, every turn | Policy: model, skill, prompt, compute per stage |
| Verification | Human taste, in the moment | Evals, judges, gates |
| What survives the run | Vibes, and maybe an edited prompt file | Lineage: artifact in/out, scores, models, costs |
| Improvements accrue to | The individual | The organisation |
| Runs on | A laptop | Cloud or on-prem |
Four dimensions decide which machine a piece of work belongs in. They are not independent — the fourth is a consequence of the first three — but they are separable enough to argue about one at a time.
Some work belongs in the interactive loop permanently. Genuinely new, large, ambiguous tasks with no verifiable definition of done are worked out by hand, in long sessions, and that is correct — there is no progression to engineer. The mistake is assuming this category is most of the work. It isn't. It is the narrow set of genuinely novel problems plus the ones whose acceptance is irreducibly a matter of taste or politics.
For everything else, there is a sharper test than frequency or size: do two or three people already do this? If several people produce acceptable results independently, a shared standard exists — even if nobody has written it down. One person repeating something might just be idiosyncrasy. Three people is a process. Plurality of practitioners proves a spec exists.
The catch is that automating it means writing that tacit methodology down, and writing it down is most of the work. The rules live in the heads of the people who do it, and they are mostly negative rules — what makes an output wrong — which nobody articulates until they see a bad one. The first artifact you need is not a prompt. It is the acceptance criteria three humans have been silently applying.
Real-world data supports the shape of this. A task-stratified study of 7,156 agent-authored pull requests found that documentation tasks were accepted at 82.1% against 66.1% for new features — and, critically, that this 16-point gap "exceeds typical inter-agent variance for most tasks." Which agent you pick matters less than what you point it at.
The two categories are not buckets, though — they are a conveyor. The fifth time you do an "ambiguous" thing, a methodology has crystallised and it has quietly become a category-two task. How fast you push items across that boundary is something you can invest in deliberately rather than wait for.
How far does the conveyor run? The most credible answer comes from a source with every commercial incentive to say "all the way." Anthropic's 2026 Agentic Coding Trends Report, citing its own Societal Impacts research, reports that while developers use AI in roughly 60% of their work, they report being able to "fully delegate" only 0–20% of tasks.
Read that gap carefully
The distance between "AI is involved in 60% of the work" and "0–20% can be handed over entirely" is not a model-capability gap that the next release closes. It is the space occupied by context, taste, verification and accountability — the four things the human in the room is supplying for free. That gap is the addressable surface of this entire article.
Two very different starting positions get confused with each other.
A new process with no history should be run manually first, in the harness. But the person doing it has a different job than it looks. They are not there to do the task. They are there to attempt the task with a harness, build the skills, push the expertise into the process, and get the success rate to something acceptable. Their unit of work is the process, not the item. Their output is not five closed tickets — it is a task type whose success rate went from 40% to 90%. That distinction is the whole discipline, and it is what separates this from "using Claude Code faster."
An established process with deep history is the opposite case, and it is common in real businesses. Here you can build the harness with the evals from the start, skip the interactive phase almost entirely, and go straight to agent-led research and optimisation — then check whether the agent reaches the quality of the current manual process.
History isn't experience. It's a labelled test set — and only if you kept it.
This reframes the dimension. The question is not how many times have we done this. It is did we retain the artifacts of having done it. A process run ten thousand times with nothing kept is exactly as cold-start as a brand new one. You cannot retrofit an eval set onto work whose inputs and outputs were never stored — which makes this a capture-time problem, and therefore a decision you are making today whether or not you know it.
Two cracks in the historical-data path, worth naming before a skeptic finds them.
The record is survivorship-biased. You have what shipped, not what was rejected and why. Negative examples are precisely what acceptance criteria are made of, and they are systematically the thing nobody keeps. If you start retaining anything today, retain the rejections.
Parity is a strange bar. Fitting to historical output inherits the manual process's defects and caps you at human quality. It is a defensible release gate — it is what a regulated buyer will ask for — but a poor goal. Gate on parity; aim past it.
This is the core of the argument, and it is about the shape of the execution, not the task.
The interactive harness is semantic-led: a continuous flow of text and code. It produces artifacts, and some of them can even be validated. What you cannot do is attribute an artifact to a specific execution. In one session you say "create me five issues," and it creates five issues. Then in the same session, "now create five PRs," and it does. Then you exit, and those PRs get reviewed by a different harness entirely. Who did what, when, caused by which instruction? There is no single history. There is no real data. It's mud.
You can run local retrospectives — look at what we learned this session, update the prompts — and that genuinely moves your setup forward. But the improvement is unevaluated and indefensible. Why did we make this change? Which failure, on which PR? Nobody can say. And an organisation that cannot tell a good change from a lucky one cannot adopt it, defend it, or roll it back. Personal productivity does not compound into institutional capability without a schema.
The decomposed workflow is data-led. Agents still do the work inside each task. The difference is that every task has edges — and edges are what measurements attach to.
Attribution requires boundaries
In one long session the unit of work has no edges, so there is nothing for a measurement to attach to. Decomposing the workflow is not a performance optimisation. It is what manufactures the join key. Artifact-in / artifact-out per task is lineage; a session is an append-only text blob with no primary key.
| Recorded per task | What it unlocks |
|---|---|
| Input artifact ID → output artifact ID | Lineage; traceability of every deliverable |
| Judge quality score | Regression analysis on success rate by stage |
| Worker model used | Cost attribution; evidence-based routing |
| Judge model used | Judge calibration; catching drift |
| Tokens consumed | Tokenomics: cost per task, per stage, per outcome |
We can show the asymmetry from our own delivery data, because we ran into it by accident. In our 90-day SDLC breakdown, the remote review workflow — decomposed into discrete review runs, each writing a row server-side — produced complete telemetry for the entire quarter: 2,214 review runs across 988 PRs, with rounds, findings, verdicts and durations per PR.
The spec and coding loops ran interactively, on a laptop. Their cost data comes from on-disk agent transcripts, and those get pruned. When we went to price the loops, only one month of the three had complete coverage. We published June's numbers rather than defend a 90-day total whose early weeks were half gone.
Same team, same quarter, same models. The only variable was where the work ran.
The granularity difference matters more than the coverage one. Server-side, the unit of record is a review run: which PR, which round, how many findings, which verdict, which model, how long. On the laptop, the unit of record is a monthly bill. You can divide it by merged PRs and get an average, which we did — about $15 per merged PR for the local loops. What you cannot do is ask which task type costs more, which stage regressed last week, or whether that prompt change helped. The number has no seams.
First: decomposition costs context, and the evidence against naive decomposition is strong. This is the part where a lot of agentic writing overclaims, so let me be precise about what the research says, because it does not say what a vendor would want it to say.
Cognition's Walden Yan argued in "Don't Build Multi-Agents" that "running multiple agents in collaboration only results in fragile systems. The decision-making ends up being too dispersed and context isn't able to be shared thoroughly enough between the agents." Anthropic's own engineering team, writing up a multi-agent research system that outperformed single-agent Claude Opus 4 by 90.2% on an internal research eval, was nonetheless explicit that coding is a poorer fit — "most coding tasks involve fewer truly parallelizable tasks than research, and LLM agents are not yet great at coordinating and delegating to other agents in real time" — and noted multi-agent systems use "about 15× more tokens than chats." A 2026 evaluation of automatically generated multi-agent systems found they "consistently underperform CoT-SC despite being up to 10x more expensive" than that single-agent baseline. And Kent Beck, after hands-on use of a coordinator/implementer/verifier system, concluded bluntly: "I've never been able to get two agents working on the same codebase at the same time without my head exploding. So I'm not convinced the swarm is the answer."
This is a real boundary, and it is not the one people think
That body of evidence is an argument against parallel multi-agent writes — several agents mutating the same codebase with dispersed decision-making. It is not an argument against a sequential workflow with boundaries and verification between stages. The two get conflated constantly, and they are different designs. Notably, Cognition's own April 2026 follow-up narrowed rather than reversed the thesis: patterns where "multiple agents contribute intelligence to a task while writes stay single-threaded" — code-review loops chief among them — do work in production. A measured workflow is that shape: one writer per stage, judges and gates around it, a durable record at every edge.
The honest version of the claim, then: measured workflows win over five hundred items, not necessarily on any single one. A long session that watched the issue get written may well produce a better PR this time. It just cannot tell you whether it usually does.
There is also a hard reason not to let any single run get arbitrarily long. Toby Ord's analysis of agent success rates proposes that performance decays as "a constant rate of failing during each minute a human would take to do the task" — an exponential, not linear, penalty for length. SlopCodeBench, which evaluates 15 coding agents across 36 problems and 196 iterative checkpoints, found no agent fully solves any problem end to end, the best passes 14.8% of checkpoints, and code quality erodes as trajectories lengthen — structural erosion rising in 77% of trajectories and verbosity in 75.5%.
Second: judge scores are not ground truth. They are model output. Run regression analysis on them long enough and you will optimise for judge-pleasing, with charts that look like data the entire way down. This is not a theoretical worry: researchers have shown LLM judges can be made to emit false positive rewards by inputs carrying no substantive reasoning at all — "non-word symbols (e.g., ':' or '.') or generic reasoning openers (e.g., 'Thought process:' or 'Let's solve this problem step by step.')" — and that the vulnerability reaches "leading proprietary systems such as GPT-o1 and Claude-4."
Our own numbers show where the ceiling sits honestly. Across 9,171 review findings with actionability computed, 39% had their cited code changed before merge; in the subset a human or agent explicitly triaged (n=1,653), 79% were judged true positives. Useful. Not gospel. This is exactly where Dimension 2 pays for itself — the retained historical set is the anchor that keeps the judge honest, which is why the dimensions are not independent.
And a gate nobody measures quietly stops existing. In that same quarter we found a contiguous block of about 20 PRs that merged with no AI review at all; review coverage was roughly 68%, not 100%. We only know because the scoreboard was instrumented.
Scaling is only really possible with workflows. So is cost management. Both need data, modelling, and operational mental models that are different from how engineers work today — closer to how you would run a factory line than how you would run an IDE.
Here is the part that makes this credible rather than hype: the units of work do not change much. A product change is still the unit of execution for a product team. The backlog survives. The org chart survives. What changes is that the team now manages the workflow that carries a change through its stages — which finally makes the old questions answerable. What is our velocity? Where is the bottleneck? What does a change cost? These are twenty years of flow questions that AI did not retire; it changed the unit prices.
Once you can answer them, an unfamiliar lever appears: the harness itself becomes the thing you optimise. An enterprise study of 22 locked evaluation tasks across six foundation models found that redesigning the orchestration layer cut tokens per task by 38% and cost per task by 41%, while raising quality per dollar by 82% — and concluded that "the orchestration layer moved cost per task more than the full spread of the model menu did."
None of tokenomics, parallelisation, A/B testing or experimentation exists without the schema. They are all management capabilities, and management capabilities need a unit of account. That is the closing move of Dimension 3, cashed out.
It is worth being clear that the payoff is real and measured, not theoretical. A study of Microsoft's early-2026 rollout of agentic CLI tools across tens of thousands of engineers found adopters "merged roughly 24% more pull requests than they would have otherwise," with the lift persisting across a four-month window. The same paper makes the cost point unprompted: "At organizational scale, token spend can run into millions of dollars annually, so misreading adoption, retention, or impact can make a rollout expensive without changing engineering velocity." Both halves of that sentence argue for the same thing.
The constraint migrates — Amdahl's law for organisations
By Dimension 1's own logic, the machine scales the routine half only. Automate it well and the bottleneck does not disappear; it moves — entirely into deciding what should be built. That is a better problem to have, but it is a different one, and it lands on a different set of people. Plan for the bottleneck you are about to create, not the one you are removing.
Schema, evals and harness are fixed costs, and you pay them per task type. Below some volume the interactive loop is simply correct and cheaper. Saying so is what separates a method from a pitch.
| Build the machine when all three hold | From | If it fails |
|---|---|---|
| Two or three people already do this task | Dimension 1 | No shared standard exists yet — stay interactive and let one crystallise |
| You retained the artifacts of having done it — or you start retaining them today | Dimension 2 | You have no test set. Start capturing now; automate later |
| Volume clears the setup cost | Dimension 4 | The fixed cost never amortises. Do it by hand and stop feeling guilty |
There is a fourth condition that overrides all three, and it comes from the constraint rather than the economics: you must be able to check the output without a human looking. METR puts a number on how demanding this gets — "some (reliability-critical and poorly verifiable) tasks require 98%+ success probabilities to be worth automating." If you cannot assert it, you cannot ship it unattended, however routine it is and however much history you kept.
"Artifacts that are hard for you to verify are often hard for users too. … Before AI, verification often happened incidentally during the process of creating work product. With AI, verification is the bottleneck."
— Hamel Husain, "'It's Hard to Eval' Is a Product Smell", June 2026
Most writing about autonomous agents spends its words on resource routing — which model, which prompt, which skill — and hand-waves verification. Routing is the cheapest dimension to automate and the one people over-index on. Verification is the wall.
And the field is currently building the wrong half of it. In a survey of 1,340 practitioners fielded in late 2025, nearly 89% had implemented observability — but only 52.4% ran offline evaluations, and just 37.3% ran online ones. Watching is nearly universal. Checking is a minority practice. Observability tells you what the machine did; evals tell you whether it was right. Only the second one lets you take your hands off.
Local execution is not merely a scaling ceiling. Work that runs on a laptop is invisible to the organisation: it cannot be audited, it cannot be parallelised, it cannot be measured, and it evaporates when the person leaves. Our own pruned transcripts are a small, cheap version of that failure — we lost two thirds of a quarter's cost history to a retention policy nobody chose. A regulated buyer asking "why did the machine do that, and on whose authority" is the expensive version.
For some systems this stops being a preference in eleven days
Under the EU AI Act, deployers of high-risk AI systems "shall keep the logs automatically generated by that high-risk AI system to the extent such logs are under their control, for a period appropriate to the intended purpose of the high-risk AI system, of at least six months" (Article 26(6); providers must design for automatic logging under Article 12). The bulk of those high-risk obligations become applicable on 2 August 2026. A six-month retention duty is not satisfiable by a folder of local transcripts that a tool prunes on its own schedule. If any of your automated work touches an Annex III use case, the substrate question has already been answered for you.
What goes in and what comes out at every step. If you cannot name them, you do not have a workflow yet — you have a habit.
Write down the methodology living in two or three heads, including the negative rules. This is the hard part and the valuable part.
Someone's job is to raise the success rate of a task type, not to close its tickets. Measure them that way.
Your current process is the benchmark; your retained history is the test set. Keep the rejections, not just the shipped work.
Not on laptops. This is the step that converts personal productivity into institutional capability.
Keep interactive loops for what they are genuinely good at — the new, the ambiguous, the taste-bound. Build measured workflows for everything else. And be honest that the second category is smaller today than the marketing suggests: the same lab that ships the most capable coding agents in the world says its engineers can fully delegate 0–20% of their work.
The question worth taking into your next planning cycle: pick the task your team does most often. Can you name its acceptance criteria, point at the last fifty inputs and outputs, and say what the success rate was last month? If not, you have not found a tooling problem. You have found the thing standing between you and any of this — and it is the same blank whether you automate the task or not.
The teams that close it will not just ship faster. They will be the only ones who can explain, with evidence, why their machine works.
How aictrl.dev helps
Disclosure: aictrl.dev builds workflow orchestration and skills governance tooling for engineering teams, so we have a commercial interest in the practices described here. The independent research cited above stands on its own, and the method works on any stack — a queue, a CI runner, and somewhere durable to write a row per task will get you most of the way. Where we help is the schema layer this article argues for: named workflow stages, a durable record per task with model, cost and verdict attached, and the skills that execute each stage kept under version control rather than in someone's shell history. If that is the layer you are missing, see how aictrl.dev can help.
Published 2026-07-22 · Analysis by Bulat at aictrl.dev, co-authored with AI