How We Evaluated AI Code Review on Our Own PRs
We ran a 30-day dogfooding assessment across 212 anonymized pull requests. The third-party reviewer had strong validity, but completeness metrics changed the decision about which review approach should be primary.
This is the companion methodology note to our AI code review tools buyer's guide. The buyer's guide compares product surfaces: PR comments, fix loops, analytics, documentation, custom rules, and public evidence. This article shows the internal assessment we would expect a technical leader to review before standardizing on a tool for code quality.
We intentionally anonymized repository names. The sample covered one platform codebase and one product application. We also avoid presenting this as a universal vendor benchmark. It is a dogfooding result from our own PRs, generated with a repeatable process and human or agent judgement where findings had to be categorized.
What this assessment does not claim
We did not compute recall, F1, or F2 because we did not build exhaustive ground truth for every possible issue in every pull request. The report measures observed valid findings, accepted or fixed findings, material-only coverage, and severity-weighted action volume across the reviewer outputs we collected.
The Deterministic Pipeline
The process is deterministic until the point where a finding needs judgement: whether it is valid, whether it was fixed, and how severe it is. Everything before and after that is structured data collection and calculation.
Ask the skill
Last 30 days, two reviewers, selected repositories.
Collect evidence
Pull requests, bot reviews, comments, and verdict sidecars.
Normalize and label
One finding schema with validity, action, severity, and category.
Calculate metrics
Valid rate, FIX count, material-only coverage, and weighted coverage.
Publish report
Pre-stage Markdown files and executive HTML summary.
Figure 1: Raw PR evidence is collected once, judgement labels are kept as auditable inputs, and all leadership metrics are calculated from the normalized dataset plus labels.
In practice, the user should not have to know the script names or artifact paths. The skill invocation can stay close to the intent:
Claude Code
/code-review-report last 30 days for my-bot and their-bot
| Artifact | Purpose | Deterministic role |
|---|---|---|
| dataset.json | Normalized GitHub PRs, reviews, bot comments, review verdict sidecars, and finding records. | Creates the same input dataset for the same repo, date, and bot filters. |
| local label JSON | Judgement labels for findings that did not already have structured verdict sidecars. | Keeps subjective triage separate from raw collection and makes it auditable. |
| analysis.json | Provider rollups, severity buckets, action counts, and coverage calculations. | Computes all metrics from the normalized dataset plus labels. |
| provider-assessment.md/html | Pre-stage Markdown and executive HTML summary. | Renders the same numbers into reviewable leadership artifacts. |
For structured verdicts, the report expects each finding to resolve into a small schema: provider, pull request, finding identifier, validity, action, severity, category, and evidence. When a review tool does not produce that sidecar, the local label files fill the same shape rather than changing the downstream calculations.
{
"provider": "aictrl.dev",
"pr": 1234,
"findingId": "finding-042",
"validity": "TRUE",
"action": "FIX",
"severity": "MAJOR",
"category": "correctness",
"evidence": "PR comment and follow-up commit"
}
The Result
The third-party reviewer looked strong if we stopped at precision. It produced 191 labelled findings, 182 of which were marked valid, for a 95.3% valid finding rate. The aictrl.dev review workflow produced 997 labelled findings, 898 of which were marked valid, for a 91.0% valid finding rate.
That was a real advantage for the third-party reviewer on trust. It also had a higher fix rate: 147 of 191 labelled findings were fixed or accepted, compared with 570 of 997 aictrl.dev findings. If the only question were "which reviewer is less noisy per comment?", the third-party reviewer would look better.
But the leadership decision was about code quality coverage. On that question, action volume and completeness mattered more. The aictrl.dev workflow found 570 fixed or accepted findings compared with 147 for the third-party reviewer. On severity-weighted fixed findings, the aictrl.dev workflow scored 2,010 versus 864.
The aictrl.dev knowledge-base enabled code review found 234 major-or-higher accepted/fixed findings across bug, high-severity, security, breaking-change, critical, and blocker classes, compared with 145 from the third-party reviewer.
Source: anonymized aictrl.dev dogfooding assessment generated by code-review-report for 2026-04-28 through 2026-05-28. Material-only excludes nits, low-severity style notes, minor issues, and consistency-only comments. Weighted coverage uses NIT=1, MINOR=2, MAJOR=5, CRITICAL=10.
Figure 2: Coverage comparison for accepted or fixed findings.
| Metric | Third-party reviewer | aictrl.dev review workflow | Decision signal |
|---|---|---|---|
| Valid finding rate | 95.3% | 91.0% | Third-party reviewer had higher per-comment trust. |
| FIX findings | 147 | 570 | aictrl.dev found 3.9x more accepted or fixed findings. |
| Material-only FIX coverage | 37.5% | 60.5% | Excluding nits narrowed the gap but did not reverse it. |
| Weighted FIX coverage | 29.4% | 68.5% | Severity weighting favored the aictrl.dev review workflow. |
Why This Changed the Decision
A low-noise reviewer is valuable. It is especially valuable when the goal is education, coaching, or keeping junior developers from being overwhelmed. But our goal for this assessment was code quality: catching correctness, reliability, maintainability, testability, and security issues before merge.
For that goal, a reviewer can be high precision and still be the wrong primary reviewer if it misses too much of the actionable issue surface. The third-party reviewer was credible. It was not useless. It simply did not produce enough of the fixed issue volume in this PR sample to be the primary quality gate.
The practical answer was not "replace commercial tools with a homegrown reviewer." The answer was to make the evaluation surface honest: count what was valid, count what was fixed, exclude nits when leadership needs material signal, and add severity weighting before making a tooling decision.
The assessment section a tech leader should ask for
Before choosing an AI code review tool, ask for an aggregate report over your own recent PRs: overlap PR count, raw and labelled finding counts, valid finding rate, accepted or fixed finding rate, material-only FIX coverage, weighted FIX coverage, representative examples, and explicit caveats about missing ground truth.
Caveats
- This is not a universal benchmark. It is an anonymized dogfooding assessment across two internal production repositories.
- This compares reviewer families, not every product in the market. The public buyer's guide handles broader product-surface comparison.
- No F1, F2, or recall claim is made. Those require an exhaustive ground-truth issue set, not just observed reviewer output.
- Labels include judgement. The pipeline is deterministic, but validity, action, severity, and category classification still require review.
- The optimization target matters. This assessment prioritizes code quality. A team focused on mentoring, onboarding, or education may weight outcomes differently.
Related Reading
- AI Code Review Tools in 2026: The PR Surface Is the Product - the companion buyer's guide for product selection.
- aictrl.dev Code Review Automation - review agents with repository context, policy instructions, and quality metrics.