AI Assisted AnalysisProduct
The AI Software Factory: Managing Change for Measurable Impact
Managed inputs, AI change workflows, and measured outcomes form a software factory. The operating model, implementation roles, and economics behind impact per cost.
Oct 1, 2026 · 17 min read · AI Assisted Analysis
ResearchProduct
We simplified our backlog screen, and users would have approved the wrong thing
Seven versions of our Backlog in two days, tested by AI agent testers and a simulated user built on Jev. Cutting 37 controls to 16 led novice testers to approve a security review instead of starting work. What the evidence caught, and where it fell short.
Sep 26, 2026 · 15 min read · AI Assisted Analysis
TutorialProduct
From a generated image to a Blender film
Build a product film from concept images, an editable 3D world and a workflow transition. The process, the Blender techniques, and the failures that taught us what to review.
Sep 9, 2026 · 18 min read · Co-authored with AI
Analysis
Tutorial
Get It Off the Laptop: From Interactive AI Loops to Measured Workflows
Anthropic's own engineers can fully delegate only 0–20% of tasks. The constraint isn't the model — it's verification, and verification needs boundaries. Four dimensions, and the rule for when to build the machine.
Jul 22, 2026 · 14 min read · AI Assisted Analysis
Research
Product
The Anatomy of an Open-Source AI Coding Agent
We mapped opencode, OpenClaw and Hermes into knowledge graphs — 28,842 files, 320,832 symbols. Despite a 6× size range their architectures rhyme: the same session, plugin, tool, provider and auth layers, plus ~25–30% platform plumbing. Then we graded them — one is markedly healthier.
Jul 19, 2026 · 9 min read · AI Assisted Analysis
Analysis
Product
AI Didn't Kill the SDLC. It Compressed It.
1,456 merged PRs and 155 production deploys in 90 days from a one-engineer team. The classic SDLC phases survived — each compressed into an AI loop with a human gate at its exit, refereed by DORA metrics.
Jul 18, 2026 · 11 min read · AI Assisted Analysis
Research
Code Review With a 12B Model: Graph Topology and the Price of Recall
A recall-first DAG pipeline around gemma-12B found real TypeScript review bugs by combining many noisy passes, graph topologies, and script checks. Workflow shape, not model smarts, did the heavy lifting.
Jun 7, 2026 · 20 min read · AI Assisted Analysis
Analysis
aictrl vs. GitHub vs. GitLab for Agentic SDLC Automation
A decision-maker guide to where GitHub, GitLab, and aictrl.dev fit across issue-to-PR automation, DevSecOps, governed skills, triggered workflows, observability, and cross-tool orchestration.
Jun 1, 2026 · 12 min read
AI Assisted Analysis
The Path to 97%: Engineering Data Platform Availability as a KPI
Your warehouse sells four nines; your data products probably run at 50-80%. A problem-source taxonomy and a weighted availability metric turn 97% into a roadmap, not a wish.
May 30, 2026 · 20 min read
Analysis
How We Evaluated AI Code Review on Our Own PRs
A dogfooding assessment over 212 anonymized pull requests showing why valid finding rate was not enough and why material and weighted FIX coverage changed the primary-reviewer decision.
May 28, 2026 · 12 min read
Analysis
AI Code Review Tools in 2026: The PR Surface Is the Product
A practical review of CodeRabbit, CodeAnt, GitHub Copilot Code Review, Qodo, Graphite, and Cursor Bugbot across PR features, inline comments, analytics, team training, docs, and eval credibility.
May 25, 2026 · 18 min read
Analysis
The Great Convergence: How Chinese AI Models Caught the West in 18 Months
DeepSeek trained a frontier model for $5.6M. Four of the top ten coding models are Chinese. Alibaba's Qwen overtook Meta's Llama in global downloads. Training costs, benchmarks, and enterprise strategy.
Apr 22, 2026 · 18 min read
AI Assisted Analysis
Proof Before Ship: How Skill Evals Turn AI Agents from Guesswork into Engineering
Over $547B in enterprise AI investment failed to deliver value in 2025. The root cause was not the models — it was shipping without proof. Here is the full eval lifecycle, data structures, and quality gate pattern.
Apr 3, 2026 · 20 min read
Research
From .feature Files to Knowledge Graphs: The Deterministic Link Between Business Intent and Working Code
68% of teams use BDD frameworks. Only 12% of stakeholders read the feature files. Knowledge graphs close the traceability gap — creating a provable chain from Gherkin scenarios to implementation code.
Mar 22, 2026 · 18 min read
Analysis
The Vibe Coding Dopamine Trap: When AI Velocity Isn't Linked to Real Outcomes
AI coding tools are faster than ever, but 56% of CEOs report zero AI ROI and AI-generated code carries 2.74x more security vulnerabilities. The problem isn't the tools — it's measuring activity instead of outcomes.
Mar 14, 2026 · 18 min read
Research
Why Knowledge Graphs Are the Missing Infrastructure Layer for Agentic AI
95% of AI pilots fail. The root cause isn't the models — it's the data layer. Research shows knowledge graphs improve LLM accuracy by 3-5x and cut token costs by 80-97%.
Feb 26, 2026 · 20 min read
Analysis
From Vibe Coding to Spec Engineering: The Discipline Gap Between Reliable and Unpredictable AI Teams
A single instruction quality change moved GPT-4 Turbo from 26% to 59% on a coding benchmark. Research shows why the best AI teams treat agent instructions as production code.
Feb 19, 2026 · 15 min read
Research
The SKILL.md Adoption Trajectory: From Convention to Standard
How a simple markdown file became the de facto standard for AI agent capabilities. We trace the adoption curve and analyze what made it succeed.
Feb 11, 2026 · 18 min read
Guide
SKILL.md Standard: The Enterprise Implementation Guide
A comprehensive guide to implementing the SKILL.md standard across your organization, from initial pilot to full rollout with governance.
Feb 7, 2026 · 22 min read
Research
The Reasoning Race: From 2.7% to 53.1% on Humanity's Last Exam
AI reasoning exploded in 12 months. Claude Opus 4.6 leads HLE at 53.1%, GPT-5.2 dominates math. Analysis of benchmark saturation, model specialization, and enterprise ROI.
Feb 5, 2026 · 10 min read
Analysis
The SaaSpocalypse: Which Software Categories AI Agents Will Replace First
Our analysis of 50+ SaaS categories reveals which are most vulnerable to AI agent disruption and the timeline for each wave of replacement.
Feb 5, 2026 · 16 min read
Research
The Swarm Paradox: Why More AI Agents Often Means Worse Results
Our research reveals that scaling AI agent teams beyond 4-5 agents introduces coordination overhead that degrades overall task performance.
Feb 5, 2026 · 12 min read
Research
The AI Electricity Crisis: Why Agent Compute Costs Will 10x by 2027
Energy consumption from AI agents is growing faster than data center capacity. We analyze the economic implications for engineering teams.
Feb 4, 2026 · 15 min read
Research
The Rise of Browser-Use Agents: From Selenium Scripts to Autonomous Navigation
Browser-use agents are evolving from simple automation to genuine reasoning about web interfaces. We chart the trajectory and key inflection points.
Feb 3, 2026 · 11 min read
Research
Why the Framework You Choose for AI Agents Actually Matters
Our comparative analysis of LangGraph, CrewAI, AutoGen, and native implementations shows surprising performance differences in production.
Feb 2, 2026 · 14 min read
Research
SWE-Bench Evolution: How Agent Performance Is Rewriting Software Engineering
An analysis of SWE-Bench performance trends and what the rapid improvement curve means for the future of software engineering teams.
Feb 2, 2026 · 13 min read