Research Product

The Anatomy of an Open-Source AI Coding Agent

We mapped three open-source AI coding agents into knowledge graphs — 28,842 files and 320,832 symbols. Despite a 6× size range their architectures rhyme; then we graded them, and one is markedly healthier than the rest.

Everyone has an opinion about how AI coding agents should be built. Fewer people have looked at how the popular open-source ones actually are. So we ran three of them — opencode, OpenClaw, and Hermes — through the aictrl knowledge-graph pipeline, which parses every file, clusters the codebase into subsystems, and rolls those up into capability domains. The result is an interactive map of each agent's architecture. Put the three maps side by side and a blueprint emerges.

What we measured

Each repo was extracted into a graph of files, functions, classes, interfaces and type aliases, then clustered bottom-up: files → subsystems (cohesive groups of ~40 files) → capability domains (the handful of top-level areas the product is organised around). Every number below is pulled live from each repo's public overview.json on the portal — nothing here is hand-estimated.

28,842
Files mapped
across 3 agents
320,832
Symbols
functions, classes, types
629
Subsystems
auto-clustered
28
Capability domains
the top-level shape
opencode's architecture rendered as an interactive C4 map on explore.aictrl.dev — 3,312 files across 11 domains.
Figure 1. What the pipeline produces: opencode's live map at explore.aictrl.dev — 3,312 files across 11 domains, rendered from the real knowledge graph. Each box is a subsystem; scroll to zoom, double-click to drill from system context down to a single function.

Code volume by agent

Figure 2. Files per repository. OpenClaw is in a different weight class — 20,915 files, more than the other two combined. Source: explore.aictrl.dev project-page counts, captured Jul 2026.

Size scales depth, not width

Here is the first surprise. OpenClaw has 6.3× more files than opencode (20,915 vs 3,312) and 6.1× more subsystems (439 vs 72). You would expect a codebase that much larger to also be organised into far more top-level areas. It isn't. OpenClaw has 10 capability domains; opencode has 11.

Across all three agents, the number of top-level domains sits in a tight band of 7 to 11 — regardless of whether the codebase is 3,000 files or 21,000. What grows with size is depth: the count of subsystems tracks file count almost linearly (72 → 118 → 439). Complexity doesn't add new top-level concerns; it deepens the existing ones.

Depth vs. width: subsystems explode, domains stay flat

Domains (width)
Subsystems (depth)

Figure 3. Domains (blue) barely move from 7–11 as the codebase grows; subsystems (magenta) climb with file count. Source: per-repo overview.json.

OpenClaw's architecture map — a dense cloud of files inside ten labelled domain boxes.
Figure 4. OpenClaw's 20,915 files at the L1 “Context” zoom. Ten domain boxes (Channel, Session, Agent, Plugin, Provider & Auth, UI…) contain a dense mesh of import/call edges — width stays small, depth fills in.

A quarter of every agent is plumbing

In every single repo, the largest domain is the same one: Platform Foundation — the pipeline's bucket for support and operational code (config, entrypoints, utilities, CI, glue) that doesn't belong to any one product capability. It is remarkably consistent: 25% to 29% of every codebase, no matter the size or language.

Share of files in the “Platform Foundation” domain

Figure 5. The undifferentiated “platform” layer as a share of each codebase. Source: Platform Foundation fileCount ÷ total files, per-repo overview.json.

Why this matters

If you're building an agent, a quarter or more of your effort will land in code that ships no user-visible capability but that everything else depends on. It's the part reviewers skim and refactors avoid — and, not coincidentally, where the graph shows the densest coupling.

They're all built from the same parts

Zoom into the subsystems and the same names keep reappearing across three independently-built projects. A session store. A plugin system. A tool layer. Model providers. An auth module. These aren't shared dependencies — each team wrote its own, in its own style — yet the pipeline, which names each cluster from its contents blind to the other repos, keeps landing on the same vocabulary. That's convergent architecture.

Building blockopencodeOpenClawHermes
Session
core-session66
sessions280
session157
Plugins
plugin120
plugins125
plugins22
Tools
tool52
openclawTools73
cli_tools106
Model providers
providers47
AIModelProviders65
providers109
Auth
auth-api44
agent-auth146
CopilotAuth46

Table 1. Three codebases, three sizes, one in a different language — each independently clusters into the same five building blocks. Each cell names the largest matching subsystem and its file count. Source: per-repo overview.json on explore.aictrl.dev.

Three teams, three codebases, no shared code — and the same five building blocks. That's the tell of a maturing category: the problem shape (drive an LLM, hold a session, expose tools, manage model providers, authenticate) is now well-understood enough that independent teams converge on it. If you're architecting a new agent, this table is a checklist.

Hermes's architecture map — labelled domain boxes for Gateway Services, CLI Tools, Platform Foundation and Search & Tools.
Figure 6. Hermes at L1. The same skeleton — a Gateway Services box (25 subsystems, 1,189 files), CLI Tools, Platform Foundation — even though, unlike the others, most of it is Python.

A TypeScript monoculture — and one Python outlier

Two of the three agents are overwhelmingly TypeScript: opencode 91%, OpenClaw 96% of files. If you're hiring for, or contributing to, this ecosystem, it's a TypeScript ecosystem. Then there's Hermes, which is 66% Python — its gateway and CLI cores are Python (gateway 99% Python, gateway_cli 100%), with a TypeScript/Electron shell bolted on top. It's the only agent of the three making a fundamentally different language bet.

Language mix by files

TypeScript Python CSS YAML Other

Figure 7. Share of files by language (100% stacked). Source: aggregated subsystem language mix, per-repo overview.json.

So what makes each one different?

If the skeleton is shared, the character is in the emphasis. The domain breakdown of each agent reads like a mission statement:

opencode — local-first & stateful

The most SQL of the three (88 files), matching its embedded SQLite engine. Clean 11-domain split: Workspace Management, LLM Orchestration, Terminal & Shell, a real Design System.

OpenClaw — a platform, not a CLI

Its biggest capability domains are Provider & Auth Services (3,031 files) and Channel (2,769). This is infrastructure — multi-tenant auth, gateways, channels — wearing an agent's clothes.

Hermes — Python & multi-model

Only 7 domains, gateway- and CLI-centric, with a distinctive Model Switch domain (437 files) — routing across models as a first-class concern, in a Python core.

We didn't just map them — we graded them

Mapping is descriptive. But the same dependency graph lets us compute the architecture-quality metrics that decades of research have tied to real defect and maintenance cost — so we ran them. Three numbers, each derived from which module depends on which:

Blast radius: how far a change ripples

Figure 8. Propagation cost — the share of each codebase a single change can reach through dependency chains. Source: computed from each repo's subsystem dependency graph on explore.aictrl.dev.

AgentEncapsulation ↑Blast radius ↓Cyclic core ↓
opencode22.9%75.3%80.6%
OpenClaw24.0%59.7%69.2%
Hermes36.0%41.5%60.2%

Table 2. Architecture-health metrics, all computed from the dependency graph (not the explorer's heuristic shadings). ↑/↓ mark whether higher or lower is better — Hermes leads on all three. Subsystem granularity; directional, best tracked within one repo over time.

One agent wins on every measure. Hermes keeps the most dependencies inside their boundaries, contains a change to ~42% of the code, and traps the fewest subsystems in the cyclic core. opencode is the cautionary tale: despite being the smallest of the three, a single change can ripple to 75% of the codebase, and 81% of its subsystems sit in one mutually-dependent knot. Small is not the same as decoupled.

Why this matters

These aren't cosmetic. MacCormack & Sturtevant (2016) showed empirically that tightly-coupled “core” code costs significantly more to maintain and carries more defects. And because the graph is exact, it doesn't just score — it can name the fix: opencode's five most out-of-place modules, reassigned to the domain their code actually talks to, would lift encapsulation from 22.9% to 39.4% — recomputed, not guessed.

The takeaway

Three teams, working independently, converged on the same shape: a session store, a plugin and tool layer, model providers, and an auth module — plus a stubborn ~25–30% of platform plumbing — organised into 7–11 top-level domains that barely move as the code grows 6×. If you're building an agent, that's your baseline blueprint. If you're evaluating one, the interesting questions aren't “does it have an LLM layer” (they all do) but how deep the platform layer runs, which language bet it made, and how tangled its dependency core is — the difference between opencode's 75% blast radius and Hermes's 42% is the difference between a codebase that fights change and one that absorbs it.

Every figure in this piece is one click away on the live maps — no login, updated from real knowledge graphs. Open a repo, zoom from system context to a single function, and check our reading against yours.

How aictrl.dev helps

Disclosure: aictrl.dev builds the knowledge-graph and workflow tooling behind these maps. The same pipeline that mapped these open-source agents runs on private repositories — turning a codebase into a navigable architecture graph for onboarding, review, and refactoring. See how aictrl.dev can help →

Sources & method

Explore the live maps at explore.aictrl.dev