From .feature Files to Knowledge Graphs: The Deterministic Link Between Business Intent and Working Code

Most engineering teams use BDD frameworks. Almost none of their stakeholders read the feature files. Knowledge graphs can close this gap — creating a queryable, provable chain from a Gherkin scenario to the function that implements it.


The Broken Promise of Living Documentation

In 2003, Dan North introduced Behavior-Driven Development with a deceptively simple premise: write specifications in plain English that both business stakeholders and machines can read. Feature files would be the bridge between intent and implementation. The documentation would stay "alive" because it was the test suite itself.

Twenty-three years later, the adoption numbers tell one story. The usage numbers tell another.

Bar chart showing BDD adoption rates vs documentation usage gap
The Documentation Gap. Most teams that adopt BDD frameworks use them for test automation, not living documentation. Stakeholder engagement with .feature files drops off sharply — only 12% of non-technical stakeholders regularly read them.

Based on our analysis of industry surveys and practitioner reports, the gap is brutal. We estimate that roughly two-thirds of engineering teams use a BDD framework — Cucumber, Behat, SpecFlow (now Reqnroll), Behave. But only about a third produce anything resembling living documentation from those frameworks. And only around one in ten stakeholders actually read the feature files that were written for them.

BDD didn't fail. It succeeded at the wrong thing. It became excellent test scaffolding and mediocre documentation. The Gherkin syntax — Given/When/Then — turned out to be a better way to structure automated tests than to communicate with product managers.

The SpecFlow Signal

When Tricentis EOL'd SpecFlow in December 2024, the .NET BDD community had to scramble. Reqnroll emerged as the community-driven successor, built directly on the SpecFlow codebase. But the disruption exposed a deeper truth: BDD tooling has been treated as disposable infrastructure, not as a strategic documentation layer. SmartBear has since transferred Cucumber stewardship to the Open Source Collective to prevent a similar fate.

Meanwhile, the documentation that stakeholders actually use — Confluence pages, Notion wikis, Google Docs — drifts further from reality with every sprint. You've seen it: a Confluence page that describes a workflow that was refactored six months ago, still confidently visited by the onboarding team.

The core problem isn't that feature files are badly written. It's that there's no structural link between a business rule described in Gherkin and the code that implements it. The link exists in developers' heads, in step definitions that glue scenarios to test code, but it's not queryable, not traversable, and not provable.

Knowledge graphs change that.


The Convergence: Three Technology Waves Meeting in 2026

Something unusual is happening in 2026. Three independent technology trends that matured separately are converging at exactly the right moment to solve the traceability problem that BDD alone couldn't.

Timeline showing convergence of BDD, Knowledge Graphs, and AI test generation
The Convergence Zone. BDD frameworks stabilized a decade ago. Knowledge graph tooling matured in 2023-2025. AI-powered test generation exploded in 2025. The intersection of all three is happening now.

Wave 1: BDD Infrastructure Is Mature and Stable

The Gherkin ecosystem is no longer experimental. Cucumber's official parser (@cucumber/gherkin) produces a well-defined AST from any .feature file. This AST is language-agnostic — the same parser works whether your step definitions are in PHP (Behat), JavaScript, Java, Python, or .NET (Reqnroll). The parsing is deterministic: given the same .feature file, you get the same structured output every time.

This matters because it means .feature files are already machine-readable structured data. They just haven't been treated that way.

Wave 2: Knowledge Graph Infrastructure Has Matured

Neo4j, Amazon Neptune, and property graph databases have moved from research curiosity to production infrastructure. a growing majority of enterprises pursuing AI are adopting knowledge graphs for context and reasoning, with analysts projecting graph infrastructure as a foundational layer for agentic development. Graph databases aren't exotic anymore — they're the standard approach when your data is fundamentally about relationships.

"An SDLC knowledge graph becomes the single source of truth connecting requirements, design, code, tests, and incidents, enabling end-to-end traceability."

— Nathan Lasnoski, "Building an Enterprise Knowledge Graph for the SDLC" (March 2026)

Wave 3: AI Can Now Generate and Validate Specifications

The third wave is the most recent. AI agents can now parse natural language requirements and generate Gherkin scenarios, and — critically — validate them against existing code. ThoughtWorks identified Specification-Driven Development (SDD) as one of 2025's key new engineering practices. GitHub launched Spec Kit, an open-source toolkit for spec-driven development. VectorCAST introduced Reqs2x, generating unit tests directly from requirements.

Forrester's Q4 2025 Wave on Autonomous Testing Platforms evaluated 15 vendors — and the category itself is new. The market has evolved from "Continuous Automation Testing" to "Autonomous Testing Platforms" that combine self-healing, adaptive, and risk-aware capabilities.

50-80%
Implementation time savings for well-specified features using SDD
aictrl estimate based on practitioner reports
$40B+
Testing automation market size in 2026, growing to $79B by 2031 at 14% CAGR
Mordor Intelligence, 2026
45%
Projected reduction in manual testing effort through AI augmentation
Industry analyst consensus, 2026

The convergence is clear: we have mature structured specifications (Gherkin), mature graph databases (Neo4j, Neptune), and AI that can bridge the gap between them. What's been missing is the architecture that connects these pieces.


.feature Files as Knowledge Graph Nodes

Here's what it looks like when you stop treating a .feature file as a test and start treating it as a first-class node in a knowledge graph.

Consider a typical feature file from an invoicing system:

Feature: Invoice Approval Workflow
  # Business rule: invoices above thresholds require manager/director approval

  Rule: Threshold-based approval routing

    Scenario: Invoice under £1,000 is auto-approved
      Given an invoice for £500
      When the invoice is submitted
      Then it is automatically approved
      And no approval task is created

    Scenario: Invoice between £1,000 and £10,000 requires manager approval
      Given an invoice for £5,000
      When the invoice is submitted
      Then an approval task is created for the submitter's manager

When you parse this with @cucumber/gherkin, you get a structured AST. When you load that AST into a knowledge graph, each element becomes a node with typed edges:

// The graph that emerges from parsing .feature files

CREATE (:Feature {name: 'Invoice Approval Workflow'})
  -[:HAS_RULE]->
(:Rule {name: 'Threshold-based approval routing'})
  -[:HAS_SCENARIO]->
(:Scenario {name: 'Invoice under £1,000 is auto-approved'})
  -[:HAS_STEP]->
(:Step {keyword: 'When', text: 'the invoice is submitted'})
  -[:IMPLEMENTED_BY]->
(:StepDefinition {pattern: '/the invoice is submitted/'})
  -[:CALLS]->
(:Function {name: 'submitInvoice', file: 'src/invoices/service.ts:42'})

Every node. Every edge. Computed, not inferred. The parser produces the feature-to-scenario-to-step chain. Static analysis produces the step-definition-to-function chain. The graph is the composition of both.

The Queryable Chain

Once the graph exists, you can ask questions that were previously impossible to answer without reading code manually:

"Which business rules does the submitInvoice function implement?" — traverse backwards from Function to Feature.

"Which scenarios have no step definition?" — find Step nodes with no IMPLEMENTED_BY edge.

"What code changed since this business rule was last verified?" — join the graph with git history.

How much of a typical codebase can a parsed KG reach? That depends on how much of your code is statically analyzable. Dynamic dispatch, reflection, and event-driven architectures create gaps that no static parser can bridge. In practice, Rath et al. (ICSE 2018) found that only 60% of commits were linked to issues even in actively maintained open-source projects — meaning 40% of the requirement-to-code links were simply missing. A parsed KG won't solve the dynamic dispatch problem, but it eliminates the human-forgets-to-link problem entirely for the edges it can compute.


The Links Already Exist — You Just Can't Query Them

Here's the key insight that reframes this entire problem: if your codebase has .feature files, the links between business rules and code already exist. They're encoded in the step definitions. A Gherkin step When the invoice is submitted is matched by a step definition with a regex pattern, which calls a function in your production code.

That's not an inference. It's a fact, encoded in your codebase right now. The problem is that this fact is scattered across files and unqueryable. You can't ask your codebase "which business rules does this function implement?" without manually reading step definitions, tracing call chains, and mapping back to scenarios.

What a knowledge graph changes

A knowledge graph doesn't discover new links. It stores the links that already exist in a structure you can query. The parser reads your .feature files and creates Feature → Scenario → Step nodes. Static analysis reads your step definitions and creates Step → StepDefinition → Function edges. The graph is the composition of both — a database of relationships that were always there but never searchable.

This is fundamentally different from the academic subfield of "trace link recovery," where researchers use information retrieval, machine learning, or LLMs to guess which requirements map to which code. Those approaches are needed when explicit links don't exist — when requirements live in Confluence and the code has no formal specification. They achieve 45–80% F1 at best (Alor et al., 2025; Hey et al., REFSQ 2025). But if you have .feature files, you don't need to guess. The links are already explicit.

The Questions a KG Can Answer

Once the links are in a graph database, you can run queries that no amount of code reading can answer efficiently:

Gap detection: "Which scenarios have no step definition?" — find Step nodes with no IMPLEMENTED_BY edge. These are business rules you wrote down but never automated.

Impact analysis: "Which business rules are affected by this PR?" — take the changed functions, traverse backwards through CALLS → StepDefinition → Step → Scenario → Feature. You know the blast radius before you merge.

Staleness detection: "Which business rules haven't been verified since the implementing code changed?" — join the graph with git history. If submitInvoice() was modified last week but the scenario that tests it hasn't run, the graph surfaces the drift.

Dead code: "Which functions are not reachable from any scenario?" — find Function nodes with no inbound CALLS edge from any StepDefinition. These are candidates for removal or missing test coverage.

Documentation generation: "Give me all business rules for the invoicing module." — traverse from a Feature node tagged "invoicing" down through Rules and Scenarios. The documentation is always current because it's derived from the same source as the tests.

Where the links don't exist

The graph can only store links that are statically computable. Step definitions use regex or pattern matching — those are parseable. Call graphs from step definitions to production functions — those are analyzable for most codebases. But some links genuinely don't exist in a form a parser can reach:

For the first case, runtime telemetry can add edges to the graph after test execution. For the second, the graph's gap detection actually helps — by showing what is specified, it makes what isn't specified conspicuous. For the third, this is where cross-system knowledge graphs become valuable, but that's a harder problem.

Rath et al. (ICSE 2018) found that only 60% of commits were linked to issues even in actively maintained open-source projects. The other 40% were simply never connected. A parsed KG doesn't solve the "nobody wrote a spec" problem. But it completely eliminates the "someone wrote a spec and nobody can find it" problem.

"Specifications become self-policing through continuous schema validation and contract testing."

— ThoughtWorks, "Spec-Driven Development: Unpacking 2025's New Engineering Practices"

Specification-Driven Development: Beyond BDD

BDD said: "Write the specification first, then implement it." SDD says: "The specification IS the system. Everything else is derived."

This is the paradigm shift happening right now. InfoQ, ThoughtWorks, and GitHub have all published frameworks for SDD in the past quarter. The core idea: .feature files (and other structured specifications) don't just describe the system — they drive code generation, test execution, documentation rendering, and deployment gates.

The Specification-Driven Development Pipeline Every edge is computed, not inferred INPUT OUTPUT .feature Files Business intent Gherkin Parser Structured AST Knowledge Graph Semantic nodes & edges Code Analysis Code refs & coverage Test Generation Executable verification Living Docs Auto-generated docs Deploy Gate Specification Transformation (KG core) Output
The SDD Pipeline. .feature files are parsed into a knowledge graph. The graph drives code analysis, test generation, living documentation, and deployment quality gates. Every edge in the chain is computed, not inferred.

Here's what the SDD pipeline looks like in practice:

1. .feature Files → Gherkin Parser → Structured AST

The @cucumber/gherkin parser is language-agnostic and deterministic. It produces Feature, Rule, Scenario, Step, and Example nodes. This is the "business intent" layer — written by product managers, QA engineers, or AI agents, in plain English.

2. Structured AST → Knowledge Graph

Each AST node becomes a graph node. Edges are typed: HAS_RULE, HAS_SCENARIO, HAS_STEP. Tags become node properties. The graph is incrementally updatable — when a .feature file changes, only its subgraph is rebuilt.

3. Knowledge Graph + Code Analysis

Static analysis of step definitions creates IMPLEMENTED_BY edges from Steps to StepDefinitions. Call graph analysis creates CALLS edges from StepDefinitions to Functions. The result is a traversable chain from business rule to implementation.

4. Derived Outputs

From the same graph, you derive:

DORA's Warning

The 2025 DORA report found that AI adoption improves individual throughput (21% more tasks, 98% more PRs) but increases delivery instability at the organizational level. More code, shipped faster, with less confidence that it does what the business intended. SDD addresses this directly: you can't ship code that doesn't have a corresponding specification edge in the graph.


What This Means for Your Team

You don't need to rewrite your codebase or adopt a new framework. If you have .feature files today — in any language, any framework — you're already halfway there.

Step 1: Parse What You Have

Run the @cucumber/gherkin parser against your existing .feature files. This works regardless of whether you use Cucumber (JS/Java), Behat (PHP), Behave (Python), or Reqnroll (.NET). The parser produces a universal AST. Load it into Neo4j or any property graph database.

Step 2: Add Code Analysis Edges

Analyze your step definition files to create IMPLEMENTED_BY edges. This is pattern matching — step definition patterns map to step text. Then trace the call graph from step definitions into your production code. These are the CALLS edges.

Step 3: Query for Insights

Even before full SDD adoption, the graph immediately surfaces:

The ROI Is Immediate

Early adopters of specification-driven approaches report 50-80% implementation time savings on well-specified features, with enterprise break-even typically within 6 months. But even the first query — "show me scenarios with no implementation" — in our experience typically reveals a significant portion of feature files that were written and never automated. That's your documentation debt, made visible.

For PHP Teams (Behat)

If your team uses Behat, the path is identical. Behat uses standard Gherkin syntax, so the same parser works. Your FeatureContext classes contain the step definitions to analyze. The knowledge graph doesn't care about the implementation language — it cares about the structural relationships between specifications and code.

For documentation generation, tools like Allure Framework (with the allure-behat adapter) can produce browsable, Confluence-like reports from your test results. But the real power comes when you move beyond reports to a queryable graph that answers structural questions about your system.

aictrl Knowledge Graph

We're building this.

aictrl's Knowledge Graph already parses TypeScript, Python, Terraform, and Cypher. Adding .feature file support means your BDD scenarios become first-class nodes — linked to the code that implements them, queryable by AI agents, and always in sync.

Join Early Access

The Deterministic Future

BDD gave us a specification language that both humans and machines can read. Knowledge graphs give us a structure that makes those specifications queryable and traversable. AI gives us the ability to generate, validate, and evolve specifications at scale.

The combination creates something that didn't exist before: a deterministic, provable link from business intent to working code. Not documentation that might be current. Not AI summaries that are probably right. A graph where every edge was computed, every relationship is verifiable, and every query returns a reproducible answer.

The teams that figure this out first won't just have better test coverage. They'll have structural confidence — the ability to answer "what does this system actually do?" at any moment, with proof.

That's not a testing improvement. That's a different way of building software.


Sources