Over $547 billion in enterprise AI investment failed to deliver value in 2025. The root cause was not the models. It was shipping without proof. Eval-driven development changes the equation — here is the framework, the data structures, and the evidence chain that turns "it looks good" into "it passes."
Software engineering solved this problem decades ago. Before CI/CD, developers shipped code based on local testing and manual review. It worked until it did not. The industry converged on automated test suites, quality gates, and reproducible builds — not because engineers lacked judgment, but because judgment alone does not scale.
AI skill authoring is where software was before continuous integration. Most teams write a SKILL.md file, try it a few times, and ship it if the outputs "look right." Forrester's 2026 testing report put it bluntly: organizations need to "test their AI agents — like, at all." The ones that adopt AI-infused testing tools report measurably higher automation coverage and fewer production defects.
The numbers tell the story. McKinsey's 2026 State of AI report found that 92% of organizations plan to increase AI investment. But only 40% have scaled beyond pilots. The gap between investment and value is not a model problem. It is a verification problem.
The eval gap in practice
Galileo AI's 2026 evaluation engineering report identified the critical issue: performance on pre-deployment tests does not reliably predict real-world utility. The "evaluation gap" between what works in a notebook and what works in production is the primary risk factor for AI deployments — and it widens when teams skip systematic evaluation entirely.
The discipline that closes this gap has a name. Practitioners call it eval-driven development (EDD) — a direct analogy to test-driven development. The core principle: prove that your skill performs inline with current metrics or better, before you accept it as a solution.
Anthropic's skill-creator plugin for Claude Code provides the most complete reference implementation of eval-driven skill development. Released in March 2026, it codifies a five-phase loop that treats every skill change as a hypothesis to be tested.
The loop works like this:
This is not just a workflow diagram. It is a negative feedback loop in the control theory sense — and recognizing it as such changes how your team thinks about AI quality.
In a negative feedback loop, the system's output is measured against a target, and the error signal (the gap between actual and desired performance) is fed back to adjust the input. The critical shift: when a skill underperforms, the fix is not to switch models or solve the problem manually. The fix is to improve the prompt, tighten the quality gates, and raise the artifact quality.
This inverts the typical failure pattern. Most teams hit a bad output and react in one of two unproductive ways: they blame the model ("Claude can't do this") or they bypass the skill and solve the problem by hand. Both responses break the feedback loop. The eval-driven approach keeps the loop closed — every failure becomes a signal that improves the next iteration of the skill itself.
What the feedback loop actually improves
Each iteration through the loop refines three things simultaneously: (1) the skill prompt — clearer instructions, better examples, more precise constraints; (2) the quality gates — sharper assertions, better-calibrated thresholds, more discriminating tests; and (3) the artifact quality — the evals.json, grading rubrics, and benchmark configurations that make the entire process reproducible. Over time, the system converges on reliably high performance — not through heroic individual effort, but through systematic error correction.
The baseline is the key
Every eval run compares two configurations: with_skill and without_skill (or old_skill for improvements). This is not optional. Without a baseline, you cannot distinguish "the skill helped" from "the model was already good at this." The delta between configurations is the actual measurement of skill value.
If this loop looks familiar, it should. It is a PDCA cycle — Deming's Plan-Do-Check-Act, the foundation of every process improvement methodology for the past 70 years. The mapping is direct:
The best engineering teams already manage software delivery this way. Google's DORA research (DevOps Research and Assessment) proved that elite teams differentiate themselves on four metrics: deployment frequency, lead time for changes, change failure rate, and mean time to recovery. The 2024 DORA report found that AI adoption improves deployment frequency — but can increase instability without strong process foundations. The eval loop provides exactly that foundation for AI skill delivery.
Six Sigma's DMAIC (Define-Measure-Analyze-Improve-Control) maps just as cleanly. The Measure phase is your baseline. The Analyze phase is your analyst pass flagging non-discriminating assertions. The Control phase is your quality gate with an acceptance threshold. These are not analogies — they are the same discipline applied to a new artifact type.
The shift: AI delivery is process management
The teams that treat AI skill authoring as a process to be managed — with metrics, feedback loops, and continuous improvement — will outperform the teams that treat it as a creative act. This is the same transition software engineering made from "artisan coding" to CI/CD pipelines. The Toyota Production System, Kaizen, PDCA, DORA — these frameworks exist because the discipline of systematic improvement always outperforms the heroics of talented individuals. Eval-driven development is how that discipline arrives in AI.
SkillsBench, the first comprehensive skill benchmark (86 tasks, 7 models), quantified the impact: structured skills improve pass rate by +16.2 percentage points on average. But the variance is enormous — from +4.5pp in software engineering tasks to +51.9pp in healthcare. This means you cannot assume a skill works. You have to measure it in your domain.
Eval-driven development requires reproducible evidence. The skill-creator framework achieves this through five JSON schemas that compose into a complete audit trail. Here is what each one captures and why it matters.
Defines the prompts, expected outputs, and verifiable expectations for each test case. This is your test harness.
{
"skill_name": "deploy-to-cloud",
"evals": [
{
"id": 1,
"prompt": "Deploy this Express app to Cloud Run with autoscaling",
"expected_output": "Dockerfile + cloudbuild.yaml + working deployment",
"expectations": [
"Output includes a multi-stage Dockerfile",
"cloudbuild.yaml has the correct project ID",
"Autoscaling is configured with min 1, max 10"
]
}
]
}
The grader produces a per-expectation verdict with evidence. Each assertion is either passed or failed, with a quote from the output proving the judgment. This is the chain of evidence.
{
"expectations": [
{
"text": "Output includes a multi-stage Dockerfile",
"passed": true,
"evidence": "Found 'FROM node:20-slim AS builder' and 'FROM node:20-slim AS runtime' in Dockerfile"
}
],
"summary": { "passed": 3, "failed": 0, "total": 3, "pass_rate": 1.0 }
}
Aggregates results across configurations and runs into mean ± stddev statistics. This is where you see the actual skill impact.
The benchmark includes a run_summary with mean and standard deviation for each metric. Multiple runs per configuration (typically 3) catch variance. The delta section shows the net impact: +50 percentage points in pass rate, at a cost of +13 seconds execution time and +1,700 tokens. That trade-off is the decision point.
For rigorous A/B comparison, a blind comparator agent scores two outputs without knowing which used the skill. It produces rubric scores across content (correctness, completeness, accuracy) and structure (organization, formatting, usability) dimensions. The winner is chosen on evidence, not labels.
Tracks the full version progression: which version beat which, the pass_rate at each step, and which version is the current best. This is the commit log of skill improvement.
Why JSON schemas matter
These are not just logging artifacts. They are machine-readable evidence that composes into an audit trail. When the EU AI Act enforcement date arrives in August 2026, organizations deploying high-risk AI systems will need documented quality management systems. Structured eval artifacts provide that documentation.
The skill-creator framework combines three distinct evaluation methodologies. Each has different strengths, and the right approach depends on what you are measuring.
| Code-Based Graders | LLM-as-Judge | Blind A/B Comparison | |
|---|---|---|---|
| Deterministic | Yes | No (variance) | No (variance) |
| Semantic coverage | Narrow | Broad | Broad |
| Cost per eval | Negligible | 1 LLM call | 2 LLM calls + comparator |
| Eliminates bias | N/A | No (position bias) | Yes (blind labels) |
| Best for | Format checks, compilation, field presence | Code quality, document structure, explanations | Proving version A beats version B |
Best for: Assertions that can be checked programmatically. Does the output contain a file with a specific format? Does the code compile? Does the API response include required fields?
The skill-creator explicitly recommends writing scripts for programmatic assertions: "scripts are faster, more reliable, and can be reused across iterations." This is the evaluation equivalent of unit tests — fast, deterministic, and cheap.
Best for: Semantic evaluation where the "right answer" is not a simple string match. Did the skill produce a well-organized document? Is the code idiomatic? Does the explanation make sense?
Recent research has converged on best practices. A 2026 study found that a 0-5 grading scale yields the highest agreement between LLM judges and human evaluators (measured by ICC). The Judge Reliability Harness, an open-source framework, provides validation suites to stress-test judge reliability across agentic and free-response scenarios. The key finding: providing both reference answers and score descriptions is critical for reliable evaluation.
"Evals skills for coding agents guard against common mistakes I've seen helping 50+ companies."
— Hamel Husain, who has helped 50+ companies and trained thousands of engineers in eval-driven developmentBest for: Deciding whether a new version is actually better. The comparator agent receives two outputs labeled "A" and "B" without knowing which used the skill. It judges quality on a rubric, produces per-dimension scores, and declares a winner with reasoning.
This eliminates position bias and label bias. The analyzer agent then reads the transcripts to identify why the winner won — producing actionable improvement suggestions with priority and expected impact.
The hybrid approach
The skill-creator does not force a choice. It composes all three: code graders for deterministic checks, LLM-as-Judge for semantic evaluation, and blind A/B comparison for rigorous version comparison. Start with code graders. Add LLM-as-Judge when you need semantic coverage. Use blind A/B when you need to prove one version beats another.
Running evals is necessary. It is not sufficient. The discipline lies in what you do with the results: establishing acceptance criteria before you run, and holding yourself to them after.
The "proof before ship" workflow has four steps:
LLM outputs are non-deterministic. A single run tells you what happened once, not what typically happens. The benchmark schema includes a runs_per_configuration field (typically 3) specifically for this reason. Mean ± stddev catches variance that a single pass/fail hides.
The analyst pass after grading specifically looks for assertions that pass 100% in both configurations. An assertion like "Output is a valid JSON file" might always pass regardless of the skill — it tests the model's baseline capability, not the skill's contribution. These assertions are noise. The analyst flags them so you can replace them with assertions that actually discriminate skill value.
BCG's RATE.ai framework validates this pattern at enterprise scale. Projects that pass their structured gate — with legal, risk, and IT sign-off built into the process — reach production 30% faster than those that try to skip ahead. The gate is not a bottleneck. It is an accelerator. And with the EU AI Act enforcement date approaching in August 2026, documented quality management systems become a regulatory requirement for high-risk AI deployments.
The cost of not evaluating
Enterprise AI development is expensive and high-stakes. Production deployments without structured evaluation routinely face significant cost overruns. Deloitte found that the vast majority of companies are still failing at AI production deployment — and the winners differ from the losers in one specific capability: systematic evaluation before deployment.
Eval-driven development is not just a best practice. It is becoming an industry. Q1 2026 saw significant market activity: Braintrust raised $80M at an $800M valuation, Langfuse was acquired by ClickHouse, and Humanloop joined Anthropic. The signal is clear: eval infrastructure is where the money flows.
The ecosystem segments into four categories:
"Eval-driven development is not optional. The era of shipping AI features built on prompt-engineering-by-anecdote is ending."
— Vercel Engineering Team, on their shift to systematic evaluation for AI features (2026)Hamel Husain and Shreya Shankar — the most visible evangelists for eval-driven development — have trained over 2,000 PMs and engineers, including teams at OpenAI and Anthropic. Their O'Reilly book on AI evaluations is expected later in 2026. The trajectory is clear: eval expertise is becoming a core competency, not a nice-to-have.
You do not need a $800M platform to start. The skill-creator framework runs locally. Here are five concrete steps to implement eval-driven development today:
evals/evals.json. Do not overthink this step. Three good prompts beat twenty weak ones.The bottom line
Eval-driven development is the difference between "I think this works" and "I can prove this works." In a world where 80% of enterprise AI investment fails to deliver value, proof is not pedantry. It is survival. The teams that build evaluation into their workflow will ship better AI. The ones that do not will ship faster — into production incidents.
What is your current eval process — ship-and-pray, or proof-before-ship?
How aictrl.dev helps
Disclosure: aictrl.dev builds skills governance and workflow orchestration tooling. We have a commercial interest in the practices described in this article. That said, the research cited here is independent and the playbook above can be executed with any tooling. If you want purpose-built infrastructure for versioning, testing, and governing agent skills at scale, see how aictrl.dev can help.
Published April 3, 2026 · Analysis by Bulat at aictrl.dev