2026-06-07 · Code Review Automation

Code Review With a 12B Model: Graph Topology and the Price of Recall

Treat code review as a graph — a DAG of model and script nodes — and ask which shape finds the most bugs, and how that scales as the graph grows. We used a small, fast model as a cheap probe to map that space across many topologies and thousands of passes.

The setup was deliberately constrained: gemma-12B, a single GPU, TypeScript review tasks, and no frontier-model calls in the review loop.

The short version: on our setup, single-pass review was recall-bound. The useful strategy was not to make one pass smarter. It was to run several decorrelated passes, repeat them, sample hotter, and union the findings. Almost every model-side precision trick we tried bought precision by throwing away too many true bugs.

One framing to keep in mind throughout: this is not a state-of-the-art claim. On scale we are behind the field — one model, one small benchmark, mostly single-run. Where this work differs from most published AI-review numbers is that it is more transparent, not more capable: we report recall and publish the levers that failed. F2 is scored strictly against the current oracle: unmatched findings count as false positives unless they are independently validated and merged into the answer key.

Why We Optimized F2

For the exploration loop, F2 was the decision metric: recall counts more than precision. False positives are not free, but the first job of an automated reviewer is candidate discovery. A missed bug can ship. A bad finding can be dismissed. In the scores below, every unvalidated novel finding counts against precision; externally validated novels can be added to the oracle and then count as true positives.

That makes single-pass review a poor shape for the problem. One model call has a hard coverage ceiling. A filter that made the output look cleaner but removed half the real defects was not a win. A noisy workflow that found many more real defects often was. Under F2, recall is the game.

Review as a DAG, Not a Prompt

The mental model that made the experiments tractable was to treat review as a directed acyclic graph of artifact-passing nodes.

Model nodes call the LLM and spend the expensive budget. A model node can be a specialist lens, such as security, correctness, concurrency, validation, or edge cases. It can be a different framing, such as adversarial or incident-response review. It can also be a verifier or judge, although those mostly disappointed us.

Script nodes are deterministic and effectively free. They can compute a coverage map from earlier findings or, on a better benchmark, run tests and static analysis. Exact work should not be handed to a probabilistic model if a deterministic program can do it.

Finally, there is the terminal merge. union keeps everything after deduplication. vote also keeps everything but annotates how many independent nodes found the same issue. judge asks a model to filter the result. That terminal operator is not an implementation detail. It decides whether the workflow is trying to maximize coverage, rank findings, or prune output.

Once review is a DAG, the design space becomes explicit: model-call budget, breadth versus depth, sampling temperature, directed versus independent resampling, and terminal merge strategy.

Topology catalog

These are the named workflow shapes used in the rest of the article. Treat this section as the map: Parts 3 and 4 explain which ideas worked, using these topology names instead of redefining the graph each time. The numeric comparisons use one baseline: the full Real-bugs oracle set, where a one-node pass starts around 0.15 recall.

DiagramDescription
T0 — Single review
1 AI node
flowchart LR
  C([code]) --> R[review] --> F([findings])
  classDef code fill:#eff6ff,stroke:#3b82f6,color:#1e3a8a
  classDef lens fill:#f5f3ff,stroke:#8b5cf6,color:#5b21b6
  classDef out fill:#ecfdf5,stroke:#16a34a,color:#166534
  class C code
  class R lens
  class F out

One pooled pass on the full Real-bugs oracle: recall 0.15; F1 about 0.16; F2 about 0.15.

Hypothesis: one straightforward reviewer gives us the baseline quality floor.

What the DAG does: code goes through one model call and exits as findings. There is no decorrelation, no ranking, and no second chance to catch a missed issue.

T1 — 5-lens union
5 AI nodes
flowchart LR
  C([code]) --> S[security]
  C --> K[correctness]
  C --> N[concurrency]
  C --> V[validation]
  C --> E[edge cases]
  S --> U{{union}}
  K --> U
  N --> U
  V --> U
  E --> U
  U --> F([findings])
  classDef code fill:#eff6ff,stroke:#3b82f6,color:#1e3a8a
  classDef lens fill:#f5f3ff,stroke:#8b5cf6,color:#5b21b6
  classDef merge fill:#ecfdf5,stroke:#16a34a,color:#166534
  class C code
  class S,K,N,V,E lens
  class U merge
  class F merge

Breadth helped, mostly by raising recall; the useful move was independent discovery plus union.

Hypothesis: different defect lenses will notice different bugs, and a union terminal will preserve the extra recall.

What the DAG does: security, correctness, concurrency, validation, and edge-case reviewers run independently on the same code. The terminal node deduplicates and keeps everything.

T2 — 5 lenses resampled ×3, unioned winner
15 AI nodes
flowchart LR
  C([code]) --> P1[round 1 · 5 lenses]
  C --> P2[round 2 · 5 lenses]
  C --> P3[round 3 · 5 lenses]
  P1 --> U{{union}}
  P2 --> U
  P3 --> U
  U --> F([findings])
  classDef code fill:#eff6ff,stroke:#3b82f6,color:#1e3a8a
  classDef lens fill:#f5f3ff,stroke:#8b5cf6,color:#5b21b6
  classDef merge fill:#ecfdf5,stroke:#16a34a,color:#166534
  class C code
  class P1,P2,P3 lens
  class U merge
  class F merge

Best recall shape among the measured workflows; repeated independent passes are the mechanism behind the full-oracle curve.

Hypothesis: once the useful lens set is known, depth should beat more prompt variety: rerun the same panel, let sampling noise expose different bugs, and keep the union.

What the DAG does: three independent rounds each run the five-lens panel. Rounds do not steer each other; the final node just deduplicates and unions the discoveries.

T3 — Directed cascade ×3 rejected
15 AI nodes
flowchart LR
  C([code]) --> R1[round 1 · 5 lenses]
  R1 -->|findings| R2[round 2 · 5 lenses]
  R2 -->|findings| R3[round 3 · 5 lenses]
  R1 --> U{{union}}
  R2 --> U
  R3 --> U
  U --> F([findings])
  classDef code fill:#eff6ff,stroke:#3b82f6,color:#1e3a8a
  classDef lens fill:#f5f3ff,stroke:#8b5cf6,color:#5b21b6
  classDef merge fill:#ecfdf5,stroke:#16a34a,color:#166534
  class C code
  class R1,R2,R3 lens
  class U merge
  class F merge

It did not beat independent resampling; steering added complexity without enough new coverage.

Hypothesis: later rounds should not waste budget rediscovering what earlier rounds already found. If we pass findings forward, the next round can look elsewhere.

What the DAG does: each five-lens round receives the previous round's findings before reviewing again. The final output is still a union, so the experiment isolates whether direction helps discovery.

T4 — 5 lenses → judge filter rejected
5 AI nodes plus a judge
flowchart LR
  C([code]) --> S[5 lenses]
  S --> J[judge filter]
  J --> F([findings])
  classDef code fill:#eff6ff,stroke:#3b82f6,color:#1e3a8a
  classDef lens fill:#f5f3ff,stroke:#8b5cf6,color:#5b21b6
  classDef bad fill:#fef2f2,stroke:#ef4444,color:#991b1b
  classDef out fill:#ecfdf5,stroke:#16a34a,color:#166534
  class C code
  class S lens
  class J bad
  class F out

Filtering deleted too many true positives before the union could help.

Hypothesis: keep the recall lift from the five-lens panel, then ask a model judge to remove weak or hallucinated findings.

What the DAG does: discovery happens first, then a terminal model gate decides which findings survive. This tests whether a same-model verifier can buy precision without destroying recall.

T5 — 5 lenses + proof-obligation rejected
5 AI nodes
flowchart LR
  C([code]) --> L["5 lenses + repro / test"]
  L --> U{{union}}
  U --> F([findings])
  classDef code fill:#eff6ff,stroke:#3b82f6,color:#1e3a8a
  classDef lens fill:#f5f3ff,stroke:#8b5cf6,color:#5b21b6
  classDef merge fill:#ecfdf5,stroke:#16a34a,color:#166534
  class C code
  class L lens
  class U merge
  class F merge

The output looked cleaner, but discovery became too timid.

Hypothesis: if every finding must include a reproduction story or failing-test obligation, weak findings should disappear and precision should rise.

What the DAG does: the evidence requirement is pushed into the discovery prompt itself, then the outputs are unioned. This is not a terminal filter; it changes what the model is willing to report.

Real-bug subset; treat differences under ~0.05 F2 as noise. The important split is not "more nodes good" but "decorrelated discovery plus union good; terminal filtering bad."

What Moved the Needle

Numbers below are single-run on one 20-task benchmark unless noted; treat any difference under ~0.05 F2 as noise, not signal (see Caveats). F2 uses the strict current scorer: unmatched findings are false positives until they are independently adjudicated and merged into the oracle. The oracle was also revised mid-project, so early and late figures aren't strictly comparable.

Against the catalog above, the positive pattern is simple: add decorrelated discovery, then keep the union. T1 is the first useful jump; T2 is the same idea repeated deeply enough for stochastic coverage to show up.

IdeaTopology referenceEffect
Breadth before cleverness T0 → T1 Five independent specialist lenses lifted real-bug recall over a single review without needing any terminal judgment. The gain came from independent looks at the same code; on the full Real-bugs oracle, the one-pass baseline starts around 0.15 recall.
Depth after breadth T1 → T2 Adding distinct lenses saturated around five; repeating the proven panel and unioning did not. On the full Real-bugs oracle, the 15-node decorrelated run reached 0.71 recall, but only 0.27 F2 once unmatched novel findings counted as false positives.
Round framing T2 variant Changing the stance across repeated rounds — neutral auditor, adversary, incident responder — reached 0.71 real-bug recall on the full 20-task set. It was the clearest decorrelation axis beyond temperature, but it still wants replication.
Hotter sampling T2 variant Temperature around 1.1 raised distinct-bug discovery; cold sampling found fewer. Once the terminal operator is union, diversity is the mechanism, not a liability. This is promising, not yet fully replicated.

The winning pattern

Start at T1, deepen it into T2, decorrelate the repeats with hotter sampling and per-round framing, then union the findings — keep everything.

What Did Not Work

The negative results are easier to read once the topologies have names. T3 tried to make repeated rounds smarter; T4 and T5 tried to make outputs cleaner. None beat the simpler union-first pattern.

Intervention Topology reference What happened
Directed resampling T3 Telling later rounds what earlier ones already found so they "look elsewhere" gave no benefit — it slightly hurt. Steering a round off its likely re-finds did not surface enough genuinely new bugs to compensate.
LLM judge T4 A strict judge deleted almost everything. A softer judge still dropped recall much faster than it improved precision.
Confidence gating T4-style gate Gemma's self-reported confidence was not calibrated enough to be useful. Raising the threshold mostly removed findings, including true positives. Precision did not improve enough to offset the lost recall.
Consensus voting T2 annotation, not a filter Findings with more votes were more trustworthy, so vote count is useful for ranking and triage. But thresholding on votes gave up too many one-off true bugs. The held-out calibration picked "keep all" as the best F2 choice.
Proof obligations T5 Asking each finding to include reproduction steps and a failing test made the output cleaner, but the model became too conservative — it traded away recall.
Thinking mode T5-style conservatism The extreme version of that behavior. Very high precision in the small probe, but almost no discoveries. As a discovery pass, a failure. It might make sense as a final verifier fed specific candidates.
Hard-class prompt T1 lens variant A lens aimed directly at the hardest defects — a dedicated authorization / data-flow / cross-function prompt — scored 0 of 3 passes on the very authz bugs it was built for. Better prompting is not the fix for the hard tail.
Static linter script-node probe Semgrep with broad rulesets flagged essentially nothing across 28 files — on this code both the misses and the noise are semantic, not the syntactic patterns a linter sees. The static checks that would help need a buildable, type-resolved repo, which isolated snippets deny.

The pattern is important: model-side precision gates made the model quieter. They did not make it reliably better at separating real defects from plausible false alarms.

The Floor: What's Reachable, and What Isn't

Discovery is wildly uneven across real-impact oracle entries. If you take every individual gemma review pass we ran — about 4,000 single-lens reviews across the benchmark — and ask, for each real-impact entry, what fraction of passes surfaced it, you get a steep long tail. Two are found by at least half of all passes; 13 of 33 are found by fewer than 5%; and three are found by none at all. A single review almost never catches the tail, which is exactly why heavy resampling is the recall engine.

13 of 33 real-impact entries are found by fewer than 5% of passes, and the three at the far right by none. Those misses are why heavy resampling is the recall engine, and why recall has a ceiling.

That tail is what makes "high recall" expensive — and it is also the thing you cannot skip. Because passes are roughly independent, the chance of catching a given entry at least once after k passes is 1−(1−pi)k, where pi is that entry's own per-pass find-rate; mean recall is the average of that over all the real-impact entries. The distinction is the whole game: if you plugged in a single average p for every entry, the curve would shoot to ~100% within about fifteen passes. The real curve crawls instead, because a handful of entries have pi near zero — the heavy tail, not the average, governs the shape. We populate pi from the measured find-rates in the chart above and check the result two ways: the analytic line, and a bootstrap that resamples our real passes (the points). They agree to within a percentage point at every overlapping budget — but only because the curve is built from the measured per-entry tail. An average alone (a single run) would badly overstate it.

Real-bugs oracle · 33 entries

The price of recall

One pass catches ~15% of real bugs. Getting to 90% takes about 500 independent passes.

15%
single-pass recall; F1 ~0.16 and F2 ~0.15
71%
15-node decorrelated run, unioned on the full 20-task oracle
91%
hard ceiling: 3 of 33 real bugs were never found by gemma
Recall and F2 vs number of reviews per task, estimated from the full pooled pass distribution. Purple = recall, grey = F2; shaded bands are 95% bootstrap intervals. Solid dots are observed topology jobs; gold triangles are single-pass cloud-model recall baselines. F2 peaks around ~0.28 and then falls because extra passes add unmatched false positives faster than they add new true positives.

So the ceiling isn't really about discovery budget — it's a capability wall. And the tail has a shape. The chronically-missed bugs are authorization/scoping, cross-function invariants, and data-flow — defects you only see by tracing values and permissions across the code. They aren't missed for want of a better prompt (we tried that and it scored zero on them); they're missed because a single-file, single-pass read can't trace them. The lever that would reach them is execution and cross-file analysis — running the tests, resolving the types — which the isolated-snippet benchmark can't host. That's the most consequential limitation, and the clearest direction for building this for real.

Relation to prior work

The curve 1−(1−p)k connects this experiment to a familiar family of sampling results. It is the pass@k metric from code-generation evaluation, written per bug; averaged over a heavy-tailed distribution of per-item rates it produces exactly the power-law approach to a ceiling that Schaeffer et al. derive for repeated LLM sampling, and the union-of-samples coverage gain is the repeated-sampling result.

But the most striking precedent is older and human. Late-1980s and 1990s software-inspection research built exactly this model — capture–recapture with a per-defect, per-reviewer detection probability — to answer a practical question about human review teams: how many reviewers do you need to find what fraction of the defects? Their answer was that defects have wildly different detection probabilities, that independent reviewers find different ones, and that pooling more of them raises coverage with diminishing returns. That is, almost verbatim, what we observe for repeated LLM passes. The useful way to read our result is therefore not "a new formula for LLMs" but "a stochastic LLM pass behaves like an independent human inspector" — so the decades-old inspection economics (decorrelate your reviewers, add more of them, expect a long tail and a ceiling) transfer to a panel of model calls. Mathematically the whole thing is also an incidence-based species-accumulation (rarefaction) curve whose asymptote is a coverage-deficit problem — the move STADS made for fuzzing.

What we add is the instantiation for small-model code review: measuring the per-bug detection-rate tail, showing it forces a hard recall ceiling, and turning it into a price-of-recall curve. The closest published analog, SWR-Bench, also runs review several times and reports recall rising with diminishing returns — but it aggregates with another model rather than a plain union, and fits no detection model, so it has neither the tail nor the ceiling nor the cost curve.

How to Read This Against the Vendor Numbers

A warning that cuts both ways. Most published AI-review numbers are precision- or "noise"-focused and rarely report recall at all; several are self-run and don't reproduce (one vendor's self-reported 82% "catch rate" came out at 45% when a third party re-ran the same repositories with a corrected dataset); and many use a "did the developer change the code" proxy instead of a defined oracle. We made the opposite choices — report recall, split the oracle into all-findings vs user-impacting, and publish the levers that failed — because those are what make a result checkable.

That is the only sense in which this is "more rigorous": more transparent, not more capable. We make no state-of-the-art claim, and on scale we are behind the field: one model, one small benchmark, mostly single-run. (For the curious, an ensemble-vote design like ours was also shipped and then retired by a major vendor at the frontier-model tier — a reason to be specific that our bet is about small local models, not a general claim.)

To compare honestly we priced everything from measured tokens, normalized to the cost of reviewing 1,000 lines of code, and drew each model's theoretical response curve from the same 1−(1−p)k model fed with its own measured per-bug and false-positive rates (gemma at budget hosting rates ~$0.05/$0.15 per 1M; GLM at its list rates). Two things stand out. First, gemma's recall curve sits above GLM-4.7's at every budget — not because gemma is the better reviewer per pass, but because its passes are ~8× cheaper, so a dollar buys far more decorrelated discovery. Second, strict F2 tells a different story: gemma's cheap resampling also emits many unmatched novel clusters, so its F2 peaks around 0.26, while GLM's smaller sample extrapolates toward ~0.37 F2 with a lower recall ceiling. So gemma owns the cost-effective recall-discovery region, not the precision region. Caveats: GLM's curve is solid only through its 7 measured passes (its ~61% ceiling is a 7-pass underestimate); resampling lifts cloud models too; and gemma carries a ~12k fixed-prompt overhead per pass that a leaner prompt would shave off.

Response curves: recall and F2 vs cost, per 1,000 LoC. Each line is a model's theoretical curve from recall(k) = meani[1−(1−pi)k], fed with its own measured per-bug and false-positive rates; the x-position is cost = per-pass cost × k (gemma $0.0011/pass, GLM $0.0092/pass, from captured tokens). Solid = measured-pass range, dashed = extrapolation. Gold stars on the F2 panel are single Haiku/Opus passes (estimated tokens). gemma rates come from ≥144 passes (reliable recall ceiling); GLM rates from only 7 passes, so its ~61% dashed recall ceiling is an underestimate. F2 counts unmatched novel clusters as false positives unless they have been independently validated into the oracle; Opus is still circular because it co-authored part of the oracle.
The public field tends to… This work chose to…
Report precision / false-positive rate (e.g. a headline "<3%"), rarely recall. Report recall and strict F2 front and centre, because single-pass review is recall-bound and precision depends on adjudication policy.
Use injected bugs, self-curated comment lists, or a "did the dev act on it" usefulness proxy as ground truth. Define the oracle and split it into full vs user-impacting (real-bug) sets, with FALSE entries and unmatched findings both costing precision until a novel is independently validated into the answer key.
Publish only what worked, on self-run benchmarks that don't always reproduce. Publish negative results — judge, confidence, consensus-vote, proof-obligation, and thinking all failed to move F2.
Run frontier cloud models. Use a small ~12B open-weight model — we found no public work demonstrating a small, self-hostable model doing competitive detection review.

Caveats

The numbers are not universal. Much of the exploration was single-rep, and the notes call out roughly plus-or-minus 0.05 noise. Small differences are not discoveries.

This was one small model on one benchmark. The benchmark used isolated snippets, so we could not run the project, run real tests, or apply static analysis with resolved imports. That blocks the most promising precision nodes.

The oracle also changed during the work. The current oracle has 129 TRUE entries, 245 FALSE entries, and 3 UNCERTAIN entries; 33 TRUE entries are tagged real/actionable. Of those 33, 27 are FIX and 6 are DEFER; the score tables half-penalize missed DEFER entries, so metric denominators can differ from the raw entry count. The team split theoretical findings from real user-impacting bugs and added confirmed false labels for hallucinated findings. Unmatched findings count as false positives unless they are independently reviewed and merged into the oracle. Part of the oracle was authored by a strong model that also appears in our comparisons, which can inflate precision for that model — another reason recall is the cleaner claim. Early and late numbers need care.

Takeaways for Engineers

If you are building multi-pass code review, do not start by writing one heroic prompt. Start with a graph. Spend model budget on decorrelated discovery: multiple lenses, repeated passes, hotter sampling, and a union terminal. Use vote counts for ranking, not filtering, if recall is your objective. Move exact work into script nodes. If you have a buildable repo, run tests and static analysis instead of asking the model to simulate them.

Most importantly, choose the metric that matches the product. If the tool must avoid annoying developers, precision-heavy gates may be appropriate. If the tool's job is to catch defects before they ship, optimize recall first, then compute F2 only with a scorer that penalizes every unreviewed novel finding or sends it through independent adjudication. A small model is best used as a high-recall candidate generator.

There is also a class of users for whom this is not a cost trade-off at all. Organisations that cannot send source code to an external model — regulated industries, defence, anything under strict data-residency or IP-isolation rules — have largely been shut out of the cloud-reviewer market. For them the comparison is not gemma-DAG versus Opus; it is gemma-DAG versus nothing. The result here is that a recall-first panel on a single local GPU, running fully on-prem or air-gapped, makes automated detection-review viable where compliance forbids the cloud — and the levers in this post (decorrelated passes, union, strict F2) are precisely how you wring useful recall out of the smaller model you are restricted to while measuring the false-positive cost. The honest framing from the frontier still holds: when you can use a frontier model it is the more efficient bug-finder; this is for when you can't.

Where this goes next: ensembles of models, and roles

Adding passes to a single model saturates — recall approaches its ceiling and F2 falls once unmatched false positives accumulate, because every extra pass redraws from the same distribution. The decorrelation that keeps paying comes from different reviewers, and the strongest form we did not test is a heterogeneous ensemble of models: distinct model families find distinct bugs, exactly as distinct human inspectors do — the reviewer-variation effect both the human and the LLM literature support. Once the panel spans models, cost optimisation stops being "how many passes" and becomes a question of DAG shape and role assignment. That is where the roles get interesting: cheap models doing high-recall discovery, and a frontier model in a judge or verifier role applied only to the surfaced candidates to cut the false positives that every within-model precision lever here failed to remove. Our own judge nodes failed because they were the same 12B doing the filtering; a stronger, different model judging a small candidate set — cheap to run, since it sees a handful of findings, not the whole file — is the experiment we would run next. The continuation of this work is not more of one reviewer; it is a well-shaped team of unlike ones.

The practical result

Gemma-12B was not a clean reviewer. It was a cheap, parallel, noisy bug candidate generator. With the right DAG, that was enough to find real defects.

Related reading

For the buyer's-guide companion to this experiment — how the commercial AI code review tools structure their PR surface, analytics, and eval claims — read AI Code Review Tools in 2026: The PR Surface Is the Product.


Updated 2026-06-09

We reran the article charts with a stricter precision policy. Previously, findings that did not match the answer key were kept as novel and did not hurt precision. The current scorer counts every unmatched finding as a false positive unless it is independently validated and merged into the golden set; once validated, that same finding becomes a normal true-positive oracle entry. We also audited the current answer key for duplicate/conflicting rows before rerunning the calculations.

The effect is material for F2, but not for recall. The recall story is unchanged: decorrelated passes still raise the 15-node run to about 0.71 recall. The F2 story is stricter: the empirical F2 curve now peaks around 0.28 instead of flattening near 0.39–0.40, and the observed 15-node decorrelated run moves from about 0.43 F2 to 0.27 F2. That is why the article now frames gemma-12B as a high-recall candidate generator rather than a high-precision reviewer.


Sources and Further Reading

Published 2026-06-07 · Analysis by Bulat at aictrl.dev