The Anatomy of an Open-Source AI Coding Agent
We mapped three open-source AI coding agents into knowledge graphs — 28,842 files and 320,832 symbols. Despite a 6× size range their architectures rhyme; then we graded them, and one is markedly healthier than the rest.
Everyone has an opinion about how AI coding agents should be built. Fewer people have looked at how the popular open-source ones actually are. So we ran three of them — opencode, OpenClaw, and Hermes — through the aictrl knowledge-graph pipeline, which parses every file, clusters the codebase into subsystems, and rolls those up into capability domains. The result is an interactive map of each agent's architecture. Put the three maps side by side and a blueprint emerges.
What we measured
Each repo was extracted into a graph of files, functions, classes, interfaces and type aliases, then
clustered bottom-up: files → subsystems (cohesive groups of ~40 files) →
capability domains (the handful of top-level areas the product is organised around).
Every number below is pulled live from each repo's public
overview.json on the portal — nothing here is hand-estimated.
Code volume by agent
Figure 2. Files per repository. OpenClaw is in a different weight class — 20,915 files, more than the other two combined. Source: explore.aictrl.dev project-page counts, captured Jul 2026.
Size scales depth, not width
Here is the first surprise. OpenClaw has 6.3× more files than opencode (20,915 vs 3,312) and 6.1× more subsystems (439 vs 72). You would expect a codebase that much larger to also be organised into far more top-level areas. It isn't. OpenClaw has 10 capability domains; opencode has 11.
Across all three agents, the number of top-level domains sits in a tight band of 7 to 11 — regardless of whether the codebase is 3,000 files or 21,000. What grows with size is depth: the count of subsystems tracks file count almost linearly (72 → 118 → 439). Complexity doesn't add new top-level concerns; it deepens the existing ones.
Depth vs. width: subsystems explode, domains stay flat
Figure 3. Domains (blue) barely move from 7–11 as the codebase grows;
subsystems (magenta) climb with file count. Source: per-repo overview.json.
A quarter of every agent is plumbing
In every single repo, the largest domain is the same one: Platform Foundation — the pipeline's bucket for support and operational code (config, entrypoints, utilities, CI, glue) that doesn't belong to any one product capability. It is remarkably consistent: 25% to 29% of every codebase, no matter the size or language.
Share of files in the “Platform Foundation” domain
Figure 5. The undifferentiated “platform” layer as a share of each codebase.
Source: Platform Foundation fileCount ÷ total files, per-repo overview.json.
If you're building an agent, a quarter or more of your effort will land in code that ships no user-visible capability but that everything else depends on. It's the part reviewers skim and refactors avoid — and, not coincidentally, where the graph shows the densest coupling.
They're all built from the same parts
Zoom into the subsystems and the same names keep reappearing across three independently-built projects. A session store. A plugin system. A tool layer. Model providers. An auth module. These aren't shared dependencies — each team wrote its own, in its own style — yet the pipeline, which names each cluster from its contents blind to the other repos, keeps landing on the same vocabulary. That's convergent architecture.
| Building block | opencode | OpenClaw | Hermes |
|---|---|---|---|
| Session | core-session66 |
sessions280 |
session157 |
| Plugins | plugin120 |
plugins125 |
plugins22 |
| Tools | tool52 |
openclawTools73 |
cli_tools106 |
| Model providers | providers47 |
AIModelProviders65 |
providers109 |
| Auth | auth-api44 |
agent-auth146 |
CopilotAuth46 |
Table 1. Three codebases, three sizes, one in a different language — each independently clusters into the same five building blocks. Each cell names the largest matching subsystem and its file count. Source: per-repo overview.json on explore.aictrl.dev.
Three teams, three codebases, no shared code — and the same five building blocks. That's the tell of a maturing category: the problem shape (drive an LLM, hold a session, expose tools, manage model providers, authenticate) is now well-understood enough that independent teams converge on it. If you're architecting a new agent, this table is a checklist.
A TypeScript monoculture — and one Python outlier
Two of the three agents are overwhelmingly TypeScript: opencode 91%, OpenClaw 96%
of files. If you're hiring for, or contributing to, this ecosystem, it's a TypeScript ecosystem. Then
there's Hermes, which is 66% Python — its gateway and CLI cores are
Python (gateway 99% Python, gateway_cli 100%), with a TypeScript/Electron shell
bolted on top. It's the only agent of the three making a fundamentally different language bet.
Language mix by files
Figure 7. Share of files by language (100% stacked). Source: aggregated subsystem
language mix, per-repo overview.json.
So what makes each one different?
If the skeleton is shared, the character is in the emphasis. The domain breakdown of each agent reads like a mission statement:
opencode — local-first & stateful
The most SQL of the three (88 files), matching its embedded SQLite engine. Clean 11-domain split: Workspace Management, LLM Orchestration, Terminal & Shell, a real Design System.
OpenClaw — a platform, not a CLI
Its biggest capability domains are Provider & Auth Services (3,031 files) and Channel (2,769). This is infrastructure — multi-tenant auth, gateways, channels — wearing an agent's clothes.
Hermes — Python & multi-model
Only 7 domains, gateway- and CLI-centric, with a distinctive Model Switch domain (437 files) — routing across models as a first-class concern, in a Python core.
We didn't just map them — we graded them
Mapping is descriptive. But the same dependency graph lets us compute the architecture-quality metrics that decades of research have tied to real defect and maintenance cost — so we ran them. Three numbers, each derived from which module depends on which:
- Encapsulation — how much of a module's dependency traffic stays inside its own domain. Higher = cleaner boundaries.
- Blast radius (propagation cost) — how much of the codebase a single change can reach through dependency chains. Lower = safer to change.
- Cyclic core — the share of the system trapped in one mutually-dependent knot. Lower = less tangled.
Blast radius: how far a change ripples
Figure 8. Propagation cost — the share of each codebase a single change can reach through dependency chains. Source: computed from each repo's subsystem dependency graph on explore.aictrl.dev.
| Agent | Encapsulation ↑ | Blast radius ↓ | Cyclic core ↓ |
|---|---|---|---|
| opencode | 22.9% | 75.3% | 80.6% |
| OpenClaw | 24.0% | 59.7% | 69.2% |
| Hermes | 36.0% | 41.5% | 60.2% |
Table 2. Architecture-health metrics, all computed from the dependency graph (not the explorer's heuristic shadings). ↑/↓ mark whether higher or lower is better — Hermes leads on all three. Subsystem granularity; directional, best tracked within one repo over time.
One agent wins on every measure. Hermes keeps the most dependencies inside their boundaries, contains a change to ~42% of the code, and traps the fewest subsystems in the cyclic core. opencode is the cautionary tale: despite being the smallest of the three, a single change can ripple to 75% of the codebase, and 81% of its subsystems sit in one mutually-dependent knot. Small is not the same as decoupled.
These aren't cosmetic. MacCormack & Sturtevant (2016) showed empirically that tightly-coupled “core” code costs significantly more to maintain and carries more defects. And because the graph is exact, it doesn't just score — it can name the fix: opencode's five most out-of-place modules, reassigned to the domain their code actually talks to, would lift encapsulation from 22.9% to 39.4% — recomputed, not guessed.
The takeaway
Three teams, working independently, converged on the same shape: a session store, a plugin and tool layer, model providers, and an auth module — plus a stubborn ~25–30% of platform plumbing — organised into 7–11 top-level domains that barely move as the code grows 6×. If you're building an agent, that's your baseline blueprint. If you're evaluating one, the interesting questions aren't “does it have an LLM layer” (they all do) but how deep the platform layer runs, which language bet it made, and how tangled its dependency core is — the difference between opencode's 75% blast radius and Hermes's 42% is the difference between a codebase that fights change and one that absorbs it.
Every figure in this piece is one click away on the live maps — no login, updated from real knowledge graphs. Open a repo, zoom from system context to a single function, and check our reading against yours.
Disclosure: aictrl.dev builds the knowledge-graph and workflow tooling behind these maps. The same pipeline that mapped these open-source agents runs on private repositories — turning a codebase into a navigable architecture graph for onboarding, review, and refactoring. See how aictrl.dev can help →
Sources & method
- Structural data: per-repo
overview.jsonfrom explore.aictrl.dev (domains, subsystems, file counts, language mix) · file/symbol totals from each project page. Captured 19 Jul 2026. - Repositories mapped: anomalyco/opencode, openclaw/openclaw, NousResearch/hermes-agent.
- Architecture metrics (encapsulation, propagation cost / blast radius, cyclic core) computed from each repo's subsystem dependency graph. Method and the empirical link to defect / maintenance cost: MacCormack, Rusnak & Baldwin (2006), Management Science 52(7) and MacCormack & Sturtevant (2016), Journal of Systems and Software 120.
- Caveats. Subsystem and domain names are generated by the pipeline (an LLM labels each file cluster), so they're descriptive, not official. Symbol counts aren't directly comparable across languages — TypeScript contributes interfaces and type aliases that Python has no equivalent for. The architecture metrics are computed at subsystem granularity and are directional — most trustworthy within one repo over time, not as a cross-repo leaderboard. The explorer's “risk” and “coverage” shadings are separate heuristic proxies and are deliberately excluded from the numbers above.