AI coding tools are faster than ever. But AI-generated code carries 2.74x more security vulnerabilities, and most teams still measure lines generated instead of incidents prevented. The real question isn't whether AI helps — it's whether anyone is measuring the right things.
By Bulat Yapparov · March 14, 2026 · 18 min read
AI coding tools are genuinely faster than they were a year ago. The speed improvements are real, the code quality has tightened, and adoption is nearly universal — 92% of developers now use them daily.
What hasn't kept up is measurement. Most teams track activity — lines generated, suggestions accepted, tasks completed — and report impressive numbers. But when researchers look at organizational outcomes, the story changes.
Faros AI's analysis of 40,000 developers found that individual task completion improved 20-40%, but organizational DORA metrics — lead time, deployment frequency, change failure rate — remained flat or worsened. The code got written faster. Nothing else improved.
Across multiple studies, self-reported productivity consistently overshoots measured productivity by 10-40 percentage points. Developers feel faster. The system-level data doesn't agree.
The tools aren't the problem. The problem is that feeling productive and being productive are neurologically different experiences — and AI coding tools are uniquely good at triggering the former without guaranteeing the latter.
The Dependency Signal
METR's February 2026 update revealed that 30-50% of developers now refuse to participate in studies if they can't use AI. This could reflect rational tool preference — you wouldn't do carpentry without power tools either. But when developers can't even evaluate their own productivity without the tool, that's worth paying attention to.
The question isn't whether your developers use AI coding tools — 92% already do. The question is whether anyone on your team is measuring outcomes, not just output.
Here's how coding used to work: you struggled with a problem for minutes or hours, tried approaches that failed, debugged your mental model, and eventually — when something finally worked — your brain released a burst of dopamine. One big reward for sustained effort. That's how skills form.
Here's how vibe coding works: you type a prompt, wait 30 seconds, and get working code. Prompt, reward. Prompt, reward. Every few seconds, another hit of dopamine. The effort-to-reward ratio collapsed to nearly zero.
"Before AI, programming gave two dopamine hits: figuring things out AND getting them to work. Now, the AI does all the figuring out. You only get the shallow reward."
— Mark Craddock, "The Vibe Code Addiction"Behavioral psychologists call this variable-ratio reinforcement — the same mechanism behind slot machines and social media feeds. The AI doesn't always produce perfect code; sometimes it nails it, sometimes it hallucinates. That unpredictability is what makes the dopamine response so powerful. B.F. Skinner documented this in the 1950s: uncertainty heightens the reward.
A January 2026 paper in the British Journal of Psychiatry described "algorithmic dopamine economies" — digital systems that decouple reward from effort. The finding that matters for engineering leaders: chronic exposure to instant digital rewards shifts behavior toward low-effort alternatives and weakens the capacity for long-horizon planning.
Andrej Karpathy coined "vibe coding" in February 2025 as a throwaway shower thought. One year later, he publicly retired the term, replacing it with "agentic engineering" — "engineering to emphasize that there is an art & science and expertise to it." The creator of the meme recognized it had become a trap.
ThePrimeagen's Reversal
After three months of sustained vibe coding on stream, ThePrimeagen — one of the most influential developer educators — publicly burned out: "I am tired of vibing on stream. I have this growing sickness that is just eating me alive around vibing." He concluded that vibe coding works only for tools you have "no desire to build yourself." Not for real engineering.
A January 2026 study from arXiv catalogued the "AI Genie phenomenon" — AI systems that give users exactly what they ask for with minimal friction. The researchers found that the AI's agreeableness deepens the dependency loop. It never pushes back. It never says "that's a bad idea." It validates every prompt, reinforcing the cycle.
The dopamine loop isn't just a cognitive curiosity — it has a measurable engineering cost. CodeRabbit's analysis of AI-generated pull requests quantified what engineering managers have been sensing: the code gets written faster, but it's significantly worse. (Note: CodeRabbit sells AI code review tools, so they have a commercial interest in this finding — but the numbers align with independent data from SonarSource and Stack Overflow.)
Stack Overflow's developer trust data tells an uncomfortable story. Trust in AI code accuracy dropped from 43% to 33% in a single year — yet adoption kept climbing. Developers trust AI tools less than they did last year, and they use them more than ever. That's not rational tool adoption. That's compulsive use.
The downstream cost lands on code review. Faros AI's data shows PR review time increased 91% in teams with high AI adoption. The bottleneck didn't disappear — it migrated upstream. AI accelerated the writing phase and created a traffic jam in the review phase.
Anthropic — the company that builds Claude — published a study of 52 junior developers learning a new Python library. The developers using AI for code generation scored below 40% on comprehension tests. Those who used AI only to ask conceptual questions scored above 65%.
"AI use impairs conceptual understanding, code reading, and debugging abilities, without delivering significant efficiency gains on average."
— Anthropic Research, "How AI Assistance Impacts the Formation of Coding Skills" (2026)Addy Osmani named the resulting pattern: the "perpetual junior." These are developers who appear productive on surface metrics — PRs merged, lines written, tickets closed — but who cannot debug independently, cannot reason about system architecture, and cannot function when the AI is unavailable. They never build the foundational skills because the AI removed the struggle that forms them.
This isn't abstract. Gartner predicts 50% of organizations will require "AI-free" skills assessments by late 2026 because they can no longer distinguish real competence from AI-augmented performance in interviews and daily work.
Important caveat: not all coding is equal
AI tools are genuinely excellent for boilerplate, test generation, migrations, documentation, and learning unfamiliar codebases. The quality tax is concentrated in novel logic, complex state management, and architectural decisions — precisely the work where the dopamine loop is most seductive and the consequences of unchecked output are most severe. The risk isn't in using AI tools. It's in using them undifferentiated across all work types.
Cursor's president said it plainly: "There is no single definitive metric for measuring the economic impact of AI on software engineering." That's unusually candid from a vendor — and it explains why every vendor dashboard shows you activity metrics instead of outcomes.
The gap matters because it shapes purchasing decisions. PwC's 2026 Global CEO Survey found that 56% of CEOs report getting "nothing out of" their AI investments — across all AI, not just coding tools. Only 12% said AI both grew revenues and reduced costs. A Danish study of 25,000 workers across 7,000 workplaces found only 3% productivity improvement from general AI tool adoption.
The enterprise response is predictable: Forrester reports companies are deferring 25% of planned AI spend to 2027. Gartner places generative AI in the "Trough of Disillusionment" on its 2026 Hype Cycle. The pressure to show real ROI is intensifying — and teams that can't distinguish activity from outcomes will lose their tooling budgets.
The companies getting real value from AI coding tools share one trait: they measure outcomes, not activity. BCG's research shows 70% of successful AI implementation budgets go to people and process changes, not tools. The tool is the easy part. The discipline is what matters.
Based on IBM's enterprise measurement framework and large-scale developer productivity research, here's what to track instead of vanity metrics:
1. AI-Attributed Incident Rate
Track incidents (bugs, outages, security events) where root cause analysis traces back to AI-generated code. If this number is rising, your AI tools are creating more problems than they solve. Measure weekly.
2. Time-to-Resolution (MTTR)
AI-generated code that no one understands takes longer to debug. If MTTR is increasing alongside AI tool adoption, your team is trading write-time savings for debug-time costs. The net may be negative.
3. Code Review Confidence Score
Survey reviewers: "How confident are you that this PR is correct?" If confidence is dropping while PR volume increases, you're flooding the review pipeline with code that reviewers can't adequately verify.
4. Deployment Frequency (DORA)
The ultimate system-level metric. If your team writes code faster but deploys at the same rate, the bottleneck moved — it didn't disappear. Only track this alongside change failure rate.
5. Technical Debt Velocity
Measure the rate of new static analysis warnings, code duplication, and dependency issues per sprint. If this accelerates with AI adoption, you're borrowing from your future self at compound interest.
fast.ai's January 2026 analysis identified the "dark flow" state — a feeling of deep productivity that masks technical debt accumulation and skill atrophy. The antidote isn't abandoning AI tools. It's wrapping them in a discipline loop:
Design before you prompt. Know what you're building and why. Generate with AI tools. Read every line — don't accept what you can't explain. Test rigorously, including edge cases the AI won't anticipate. Review with humans who understand the system context. Measure the five outcome metrics above, not the vanity metrics your vendor dashboard shows you.
Treat AI output like code from a prolific but unreliable junior developer. Fast, eager, and wrong in ways that won't surface for weeks.
"Today, programming via LLM agents is increasingly becoming a default workflow for professionals, except with more oversight and scrutiny."
— Andrej Karpathy, February 8, 2026 — retiring "vibe coding" in favor of "agentic engineering"The vibe coding dopamine trap is real, documented, and measurable. It's also not a reason to ban AI coding tools. The teams that are winning aren't the ones avoiding AI — they're the ones that wrapped it in discipline before the dependency set in.
This week: Pull your DORA metrics from before and after AI tool adoption. Compare them. If deployment frequency and change failure rate haven't improved, you have a measurement problem — not a productivity gain.
This month: Implement AI-attributed incident tracking. Every post-mortem should ask: "Was the root cause in AI-generated code?" Build the dataset that tells you the real cost.
This quarter: Run an AI-free skills assessment to establish your team's actual capability baseline. The gap between AI-assisted performance and unassisted competence is your fragility measure.
The dopamine will keep flowing regardless. The question is whether you'll let it drive your engineering decisions, or whether you'll measure what actually matters.
Disclosure: aictrl.dev provides workflow orchestration and governance tooling for AI-assisted engineering teams. This analysis reflects our perspective on the measurement challenges facing the industry.
Former Director of Data Engineering at HeliosX, where he built the Data Platform at scale using AI. Now building aictrl.dev — the observability and governance layer for enterprise AI agents.