2026-02-19 · Research & Analysis

From Vibe Coding to Spec Engineering: The Discipline Gap Between Reliable and Unpredictable AI Teams

A single instruction quality change moved GPT-4 Turbo from 26% to 59% on a coding benchmark — zero model changes. Research from EMNLP, SkillsBench, and harness engineering converges on one finding: how you structure agent instructions matters as much as which model you choose. Here is what that means for engineering leaders.

26% → 59%
GPT-4 Turbo performance swing from a single instruction quality change — zero model changes (Can.ac, "The Harness Problem")
+16.2pp
Pass rate improvement from curated, structured skills vs unstructured instructions (SkillsBench, arXiv 2602.12670)

The $2 Billion Signal Everyone Misread

In early 2026, Meta reportedly acquired Manus — an agentic AI execution platform — for approximately $2 billion. The reflexive read was: Meta bought another AI startup. The strategic read is more interesting: Meta paid $2 billion not for a foundation model, but for agent execution infrastructure.

This distinction matters. It reflects where the AI value chain is migrating. The model layer is commoditizing. The orchestration layer — the harnesses, agents, and structured instructions that make models do predictable, reliable work — is where defensible value now accumulates.

Andrej Karpathy captured the implication in early February 2026: in his framing, you are not writing the code directly 99% of the time — you are orchestrating agents who do. He coined "agentic engineering" to describe the emergent discipline. Sam Altman put numbers to the shift: the developer time split that was 80% coding, 20% architecture is inverting to 30% coding, 70% architecture and orchestration.

This is not a prediction. It is the present state of the most AI-fluent engineering organizations. And it raises an urgent question for every engineering leader: if your developers are increasingly writing instructions rather than code, how do you apply the same rigor, tooling, and review culture to instructions that you apply to code?

The answer is emerging from both academic research and practitioner experience. The practice it validates deserves a name: Skill Engineering.

Terminology note: The evolution from prompts to skills

The field has gone through several naming cycles. "Prompt engineering" described single-turn instruction optimization. "Context engineering" (now Gartner's preferred term) describes the broader discipline of structuring all inputs to an AI system. "Skill engineering" describes the practice of treating agent instructions as versioned, tested, first-class software artifacts — the natural extension once skills become the unit of AI capability packaging.

Terminology evolution timeline from prompt engineering to skill engineering, 2022-2026
Figure 1: The terminology evolution, 2022–2026. Each naming cycle reflects a genuine shift in what developers spend their time doing, not just relabeling. "Skill Engineering" is the emerging name for treating agent instructions with the same rigor as production code. Source: aictrl editorial synthesis; Gartner context engineering report; Karpathy tweet Feb 4, 2026.

The Problem: Vibe Coding Does Not Scale

"Vibe coding" — Karpathy's term for letting AI generate code based on natural language intent without structured specification — is a legitimate starting point. It is not a production strategy. Forrester named the trajectory clearly in their 2026 predictions: "Vibe coding becomes vibe engineering," but only when engineers fail to add the specification layer that makes AI output deterministic and auditable.

The downstream consequences are measurable. Gartner predicts a 2500% increase in software defects from unstructured prompt-to-app by 2028. This prediction describes the worst case: organizations deploying AI code generation with no specification discipline at all. Any structured approach — including well-written Markdown — mitigates this risk. The question is how much structure, and at what level of formality.

The counterexample is equally clear. SkillsBench (arXiv 2602.12670), which evaluated agent skill quality across 86 tasks, found that curated, structured skills raised pass rates by +16.2 percentage points versus baseline. Crucially, self-generated skills — where the agent produces its own instructions — provided zero benefit. The improvement came entirely from human-crafted, structured skill specifications.

The "Harness Problem" research quantifies the same dynamic from a different angle. Changing a single harness variable — how an agent is instructed to operate — moved GPT-4 Turbo from 26% to 59% on a coding benchmark. Zero model changes. LangChain's evaluation on Terminal Bench 2.0 showed +13.7 percentage points from harness engineering alone. The quality of agent instructions matters as much as the quality of the model.

The "promptware crisis" is already here

Chen et al. (arXiv:2503.02400) coined the term "promptware engineering" and identified a "promptware crisis" parallel to the 1968 software crisis: uncontrolled proliferation of prompts treated as throwaway text rather than maintainable artifacts. The research explicitly calls for "prompt version control, testing frameworks, and formal specification" as the resolution. This is the Skill Engineering agenda.

Addy Osmani's position on spec-first development cuts to the practical point: architecture planning is the part vibe coders skip, and it is exactly where projects go off the rails. Olio Apps, building production systems with Claude Code, found skills unreliable when used as passive rules but highly effective when reframed as explicit process descriptions. The broader Claude Code community experience confirms it: CLAUDE.md files are suggestions the model can choose to ignore — unless they are structured in ways that make compliance unambiguous.

What This Looks Like in Practice

To understand why format matters, consider a common agent task: a deploy-pipeline agent that decides which stages of a deployment to run based on the type of change. In Markdown, the specification might look like this:

markdown Typical Markdown agent spec
## Deploy Pipeline Agent

### Stages
| Stage       | ID       | Default Runner     |
|-------------|----------|--------------------|
| Build       | build    | ci-runner          |
| Test        | test     | ci-runner          |
| Staging     | staging  | deploy-agent       |
| Production  | prod     | deploy-agent       |

### Rules
- Not every change needs all stages. Skip stages
  that are not relevant to the change type.
- Run tests before any deployment stage.
- Production requires staging to pass first.
- Ensure each stage has clear success criteria.

This reads well to a human reviewer. But for a language model, every sentence requires interpretation. What does "not relevant" mean for each stage? What counts as "clear success criteria"? The model must infer intent from prose, and different runs may infer differently.

Now consider the same specification in Python pseudocode:

python Python pseudocode agent spec
class Stage(Enum):
    """Pipeline stages in execution order. Lower ordinal = earlier execution."""
    BUILD   = ("build",   "Build & Compile",    "ci-runner")
    TEST    = ("test",    "Test Suite",         "ci-runner")
    STAGING = ("staging", "Staging Deploy",     "deploy-agent")
    PROD    = ("prod",    "Production Deploy",  "deploy-agent")

    def __init__(self, id: str, label: str, runner: str):
        self.id = id
        self.label = label
        self.runner = runner

# Execution order: each stage depends on all stages before it
STAGE_ORDER = [Stage.BUILD, Stage.TEST, Stage.STAGING, Stage.PROD]

The model does not execute this Python. It reads it as a specification. But the structural properties are immediately different: id, label, and runner are always bundled together and cannot be mismatched. STAGE_ORDER is an explicit list, not a sentence about "running tests before deployment." The model does not need to infer ordering from prose.

The practical difference is observable in two ways: reduced token consumption (the model spends less time figuring out what to do) and improved output consistency (less ambiguity means fewer divergent interpretations across runs). The EMNLP and SkillsBench research corroborates both effects at scale (see Part 5).

Why the Structure Matters: A Technical Dissection

The performance improvement is not a curiosity — it has a mechanistic explanation. Each Python construct maps to a specific cognitive property that reduces ambiguity in model interpretation.

Domain rules as match/case

The Markdown instruction "skip stages that are not relevant" is a vague directive. A Python pseudocode version replaces it with a function that defines relevance precisely:

python Pseudocode: needs_stage() with match/case
def needs_stage(change: Change, stage: Stage) -> bool:
    """Determine if this change requires this pipeline stage."""
    match stage:
        case Stage.BUILD:
            return any_file_matches(change, ["*.py", "*.ts", "*.go",
                                              "Dockerfile", "requirements.txt"])
        case Stage.TEST:
            return needs_stage(change, Stage.BUILD)
        case Stage.STAGING:
            return change.target_env in ("staging", "production")
        case Stage.PROD:
            return change.target_env == "production" and change.approved

The model now knows exactly which conditions trigger each stage. There is no ambiguity about what "not relevant" means for builds versus deployments. The test rule — "run tests if and only if there is a build" — is expressed as a function call, not as a sentence that requires interpretation.

Quality gates as executable validation

Markdown specifications typically end with a checkbox list of quality criteria. A Python pseudocode version replaces that with a validation function:

python Pseudocode: validate() quality gate
MIN_CHECKS_PER_STAGE = 3

def validate(plan: DeployPlan) -> list[str]:
    """Run before returning. All checks must pass."""
    errors = []
    for stage in plan.stages:
        if len(stage.success_criteria) < MIN_CHECKS_PER_STAGE:
            errors.append(f"{stage.label}: only {len(stage.success_criteria)} criteria, need >= 3")
        if not stage.rollback_plan:
            errors.append(f"{stage.label}: no rollback plan defined")
    return errors  # Empty list = all good

A checkbox list is a reminder to check something. A validate() function with an assert len(errors) == 0 downstream is a contract. The model understands the difference: the Python structure signals that these constraints are enforced, not suggested. Every property of the specification that reduces ambiguity contributes to more consistent agent behavior.

Before/after comparison showing dramatic performance gains from harness engineering alone on GPT-4 Turbo and LangChain agent
Figure 2: The Harness Problem — instruction quality dominates model quality. GPT-4 Turbo jumped from 26% to 59% from a single harness variable change (zero model changes). LangChain's evaluation on Terminal Bench 2.0 showed +13.7pp from harness engineering alone. Source: Can.ac "The Harness Problem" (Feb 2026); LangChain blog (Feb 2026).

The Research That Validates It

The case for structured specifications rests on two independent lines of evidence: academic research showing pseudocode outperforms prose, and industry benchmarks showing that instruction quality dominates model choice.

Academic evidence: structured instructions outperform prose

Mishra et al. (EMNLP 2023, "Prompting with Pseudo-Code Instructions") evaluated pseudocode vs natural language instructions across 132 tasks. The results: 7–16 F1-point gains for classification tasks and 12–38% ROUGE-L improvement across all tasks. The explanation aligns with the mechanistic account from Part 4: structured instructions reduce the output space the model must search, directing attention toward task-relevant reasoning paths.

Further academic work corroborates the direction. Kumar et al. (2025) found improvements across 11 instruction-following benchmarks from pseudocode fine-tuning — a training-time effect that suggests the benefit of structure is fundamental to model processing, not just a prompting trick. The FASTRIC research formalized the endpoint: a Formal Prompt Specification Language using Finite State Machines that makes LLM interactions verifiable. The DFAH paper found a moderate positive correlation (r=0.45) between instruction determinism and faithfulness.

Industry confirmation: specification dominates model choice

Can.ac's "The Harness Problem" analysis is the most striking single data point. One harness variable change moved GPT-4 Turbo from 26% to 59% on a coding benchmark — a 33 percentage point swing from instruction quality alone. LangChain's evaluation on Terminal Bench 2.0 showed +13.7pp from harness engineering. The principle: structured specifications are architectural patterns for agent behavior, and they matter as much as model selection.

SkillsBench (arXiv 2602.12670), the largest systematic evaluation of agent skill quality to date, found three directly relevant results:

  1. Curated skills raise pass rates by +16.2 percentage points versus unstructured instructions
  2. Self-generated skills provide zero benefit — the human specification layer is non-negotiable
  3. Focused skills outperform comprehensive documentation — scope discipline matters more than volume

Code-formatted reasoning: the mechanism at scale

A growing body of research shows that code-like instructions do not just improve LLM outputs — they fundamentally change how models reason. Chae et al. (EMNLP 2024, "Language Models as Compilers") introduced a THINK-AND-EXECUTE framework where models generate task-level pseudocode, then simulate its execution. Pseudocode outperformed natural language guidance for algorithmic reasoning, even though the models were trained primarily on natural language instructions.

The effect is not limited to reasoning benchmarks. Gao et al. (PAL, ICML 2023) showed that program-formatted reasoning steps outperform chain-of-thought by +8% on GSM math and +40% on GSM-hard. Li et al. (Chain of Code, 2023) achieved 84% accuracy on BIG-Bench Hard — a 12-point gain over chain-of-thought — by formatting reasoning as flexible pseudocode with an LM-augmented code emulator. Li et al. (SCoT, TOSEM 2025) demonstrated +13.79% Pass@1 improvement in code generation from structured chain-of-thought prompting using programming constructs (sequential, branch, loop).

For agent systems specifically, Wang et al. (CodeAct, ICML 2024) found that agents using executable Python code as their action space achieved up to 20% higher success rates than agents using JSON-formatted tool calls — across 17 different LLMs. The code format allows dynamic revision and multi-turn interactions that structured-JSON alternatives cannot support.

Format sensitivity: the underestimated variable

He et al. (2024) formatted identical contexts as plain text, Markdown, JSON, and YAML, then measured performance across NL reasoning, code generation, and translation tasks. GPT-3.5-turbo performance varied by up to 40% depending on the prompt template. GPT-4 was more robust but still format-sensitive; all differences were statistically significant (p<0.05). The practical implication: format is not cosmetic. It is a performance variable on par with model selection.

OpenAI's GPT-4.1 prompting guide confirms the trend from the vendor side. GPT-4.1 follows instructions "more literally" than previous models, making structured prompt formats more important than ever. The recommended structure is explicit: Role/Objective, Instructions with sub-categories, Reasoning Steps, Output Format, Examples, Context. Implicit rules are no longer being inferred — explicit structure is required.

One important caveat: Rui et al. (EMNLP 2024, "Let Me Speak Freely?") found that strict output format constraints (forced JSON, XML) can degrade reasoning by 10–15% on math and symbolic tasks. The distinction matters: structured input instructions consistently help, while overly constrained output formats can hurt. This is why specification-level pseudocode — which structures the instruction, not the output format — is the right level of abstraction.

Important nuance: structure is the proven variable, not Python specifically

The research establishes that structured specifications outperform unstructured prose. It does not directly compare Python pseudocode against other structured formats (TypeScript interfaces, YAML schemas, formal DSLs). We recommend Python pseudocode based on its combination of readability, type expressiveness, and LLM familiarity — but this is an informed recommendation, not a research finding. Teams should evaluate the format that best matches their stack and expertise. The key insight is that any well-structured format will outperform prose.

Horizontal bar chart showing pass rates by instruction quality from SkillsBench experiments
Figure 3: Pass rate by instruction quality. Self-generated skills and unstructured prose perform identically. Curated structured skills add +16.2 percentage points (SkillsBench, arXiv 2602.12670). Key insight: curated > comprehensive > self-generated ≈ none. Source: arXiv:2602.12670 (SkillsBench), February 2026.

The Industry Is Already Moving

The signal is not only in research papers. The infrastructure and community surrounding structured agent specification is maturing faster than most engineering leaders realize.

The emerging standard stack

The Agent Skills open standard — adopted by OpenAI Codex, Google ADK, GitHub Copilot, Microsoft, and Cursor — has established SKILL.md as the de facto packaging format for agent capabilities. This is the distribution layer. What lives inside those skill files is the next frontier, and Python pseudocode is a compelling answer.

Boris Cherny, the creator of Claude Code, treats his CLAUDE.md as iteratively-refined living code — a ~2,500-token document updated every time the agent does something wrong. His framing is explicit: the instruction file is code, deserving the same maintenance discipline as a critical module. This is Skill Engineering in practice from the person who built the most widely adopted agentic coding environment.

Simon Willison, whose observation that "Claude Skills are awesome, maybe a bigger deal than MCP" sparked substantial community discussion, points to the composability of structured skills as the key differentiator. A Python pseudocode skill is more composable than a prose skill because its I/O contracts are explicit.

Cursor's modular rule files

Cursor deprecated its monolithic .cursorrules format in favor of modular .mdc rule files. This is the IDE ecosystem's acknowledgment of a fact Skill Engineers already know: monolithic prose specifications do not compose. You cannot unit-test a section of a prose document. You can unit-test a typed function. The migration to modular, structured rule formats is a structural endorsement of the Python pseudocode direction.

NVIDIA shipping Cursor Rules developer guides in their enterprise toolkit is the final indicator that this is no longer an avant-garde practice — it is enterprise-grade infrastructure guidance.

The skill economy as market signal

SkillsMP crossed 160,000+ indexed skills within weeks of launch. SkillHub applies AI-evaluated quality scoring on five dimensions to 7,000+ skills. The market is discovering, in real time, that skill quality is not uniformly distributed — and that structure is the primary differentiator of high-quality skills. Python pseudocode is a natural next step: the highest-quality tier of skill specification.

The investment signal reinforces the market signal. a16z allocated $1.7 billion to AI infrastructure in 2026, with a clear thesis that agent instructions are evolving into source code — versioned, tested, and deployed with engineering rigor. Braintrust's $80M Series B for AI observability and prompt management is building the measurement layer for this stack.

"Conceptual purity matters far less than predictable behavior."

— Olio Apps, on Claude Code skills in production systems, 2026

What This Means for Engineering Leaders

The shift toward Skill Engineering has concrete team structure and hiring implications that are already being recognized by major consulting firms, even if the vocabulary has not yet converged.

The new roles

Deloitte's 2026 State of AI in the Enterprise identifies emerging AI orchestration roles across the enterprise. AI architect positions are doubling, from 30% prevalence among technology leaders to a projected 58%. Forrester describes two archetypes emerging in the post-vibe-coding organization: "Product Engineers" who interface between AI capability and product requirements, and "High-Coding Architects" who design systems specifically for AI comprehensibility.

BCG reports that leading organizations are targeting 50+ reusable agent blueprints by mid-2026. The teams maintaining those blueprints are doing Skill Engineering, even if they do not use that term. The practices — version control, peer review, structured testing, quality gates — are identical to software engineering practices, applied to instruction artifacts rather than code.

The hiring signal

Forrester's 2026 data shows CS enrollment dropping 20% and engineering hiring time doubling. These are lagging indicators of a structural shift in what the role requires. The engineer who can write precise, testable, structured agent specifications is not the same profile as the engineer who writes efficient algorithms. The overlap is high but not complete. Your hiring bar needs to account for this.

The practical question is not "do you know Python" but "can you write a Python pseudocode specification that a language model will follow predictably at production quality." This is a distinct skill, and right now almost no candidates have been explicitly trained for it. The organizations building internal skill libraries and reviewing them with the same rigor as code reviews are building this capability organically.

The EU AI Act compliance angle

EU AI Act enforcement for high-risk systems begins August 2026. The Act requires deterministic, auditable AI behavior with documented decision logic. Python pseudocode specifications are inherently more auditable than prose: every rule is explicit, every branch condition is visible, every quality gate is testable. Organizations in regulated industries should treat structured skill specifications as compliance artifacts, not just engineering conveniences. This creates a second axis of ROI beyond productivity gains.

Bar chart showing Gartner's predicted 2500% increase in software defects from unstructured AI development vs structured skill engineering staying near baseline
Figure 4: The defect divergence. Gartner predicts a 2500% increase in software defects from unstructured prompt-to-app by 2028. Structured skill engineering keeps defect rates near baseline by imposing specification discipline on AI-generated code. The gap between the two paths widens exponentially. Source: Gartner Predicts 2026; aictrl projection for structured path.

The New Stack: What Skill Engineering Looks Like in Practice

Skill Engineering is not a job title yet. It is a set of practices that high-performing AI engineering organizations are converging on independently. Here is what the stack looks like:

Layer 1: The specification format

For complex, multi-step agent tasks, we recommend a structured format with type constraints, explicit I/O, and validation logic. Python pseudocode is a strong choice because of its readability and LLM familiarity, but TypeScript interfaces, YAML with schemas, or other typed formats can achieve similar results. The key is structure, not syntax. For simpler, single-turn skills, well-structured Markdown with explicit sections and bullet-listed rules may be sufficient.

The decision rule: if your agent instruction has branching logic (do X if Y, skip Z if not W), conditional flows, or quality gates, Python pseudocode will outperform prose. If it is a simple capability description with no branching, structured Markdown is adequate.

Layer 2: The packaging standard

SKILL.md is the packaging format. Your Python pseudocode specification lives inside the skill body. The SKILL.md front matter handles discovery, metadata, permissions, and compatibility. This separation of concerns — packaging from specification — is the same principle as interface vs implementation in software engineering.

Layer 3: The governance layer

Skills should be version-controlled. Changes to skill specifications should go through peer review. Skill quality should be validated with test inputs before deployment. Breaking changes to a skill's behavior should be communicated to dependent teams, just as breaking API changes are communicated.

This is not aspirational. BCG's leading organizations already do this with their 50+ agent blueprints. The organizations not doing this are accumulating skill debt — the agentic equivalent of technical debt, manifesting as unpredictable agent behavior that is difficult to debug because the specification that produced it was never made explicit.

Layer 4: The measurement layer

OpenAI's "Testing Agent Skills with Evals" framework treats skills as testable, scorable artifacts with CI-like scoring pipelines. Braintrust raised $80M Series B on AI observability and prompt management — the measurement infrastructure for this layer. The organizations that will lead in agentic engineering are those that can answer: "What changed in this skill version, and how did it affect agent performance?" Right now, almost no organizations can answer this question.

What this costs

Structured specifications are not free. Honest costs to plan for:

An emerging practice, not yet a standard

While structured prompting is well-established (Anthropic, OpenAI, and others document it extensively), the specific practice of using typed Python pseudocode for complex agent specifications is still emerging. The teams that develop and share specification patterns — public repositories, shared templates, community benchmarks — will help shape best practices for the field. Early conventions tend to stick: this is how SKILL.md went from community practice to industry standard.

The Four-Step Skill Engineering Playbook

For engineering leaders who want to begin treating agent instructions as first-class software artifacts, the entry point is lower than it appears.

Step 1: Audit your existing instruction artifacts

Every engineering team using an agentic coding environment already has instruction artifacts — CLAUDE.md files, .cursorrules, system prompts, agent specifications. Most of these are not under version control. Most have never been peer-reviewed. Start by inventorying them, adding them to your primary repository, and creating a pull request process for changes.

Step 2: Identify your highest-value agent for conversion

Find the agent instruction that runs most frequently or whose output quality matters most to your workflow. Convert it from prose to Python pseudocode. Keep the Python structure at the specification level — you are not writing library code. You are writing a typed, structured description of process, data, and rules that a model will interpret.

Step 3: Measure before and after

Run the same benchmark inputs through both versions. Measure token consumption, output format adherence, and output quality on your task-specific rubric. The magnitude of improvement will depend on how structured your current instructions already are and how complex your agent's task is — but the research consistently shows structured specifications outperform unstructured prose.

Step 4: Establish the review culture

Skill quality degrades without the same culture that keeps code quality high: peer review, testing, and accountability for changes. The Boris Cherny model — update CLAUDE.md every time the agent does something wrong — is a minimal viable process. A mature process adds: peer review before merging, a test suite of benchmark inputs, and a changelog that captures behavioral changes over time.

The Bottom Line

The best AI engineering teams are already doing what the research supports: applying software engineering discipline to agent instruction artifacts. Skill Engineering is the emerging name for this practice. Python pseudocode is a strong format choice for complex multi-step agent tasks — its type constraints, explicit I/O, and match/case rules map directly to the properties that research from EMNLP and SkillsBench associates with better agent performance. But the deeper insight is format-independent: structured specifications outperform unstructured prose, and instruction quality matters as much as model quality.

Three actions for this quarter: first, put your instruction artifacts under version control. Second, convert your highest-frequency agent specification to a structured format — Python pseudocode for complex branching logic, well-structured Markdown for simpler capabilities — and measure the result. Third, establish a peer review process for skill changes with the same expectations you have for code changes.

The companies that build this muscle now are building organizational capability that compounds. The companies that do not are accumulating skill debt — unpredictable agent behavior that is difficult to debug because the specification that produced it was never made explicit.

How aictrl.dev helps

Disclosure: aictrl.dev builds skills governance and workflow orchestration tooling. We have a commercial interest in the practices described in this article. That said, the research cited here is independent and the playbook above can be executed with any tooling. If you want purpose-built infrastructure for versioning, testing, and governing agent skills at scale, see how aictrl.dev can help.


Sources and Further Reading

Published 2026-02-19 · Analysis by Bulat at aictrl.dev