Technology & Research · · Team LATENT
Can AI check AI-written code? What GitHub ReviewBench measures
GitHub’s ReviewBench measures AI code review through coverage and finding quality. We explain its 219 public PRs, what 96.6% agreement means, and how to interpret the results for real development.
After asking AI to write code, it is easy to ask another AI to check it. A long review does not tell you whether that check was good. Too few useful findings create extra work for the developer; too many missed issues leave defects behind.
ReviewBench, announced by GitHub on October 5, 2026, is an open benchmark for measuring those two aspects separately. It gives review agents the same code changes and examines how many known issues they find and how useful their findings are. That can inform delegation by showing what a reviewer catches and where it falls short.
Key takeaways
- ReviewBench evaluates coverage and finding quality on 219 PRs from 187 public repositories. Its sampling is informed by 103.9 million PRs, but the evaluation itself uses 219 and gives greater weight to substantial changes.
- Grounded metrics use existing labels; augmented metrics also classify findings without a reference match. Grounded precision excludes unmatched findings. Because novel findings change the denominator of augmented recall for each agent, grounded recall is the common basis for comparing coverage across agents.
- GitHub evaluated a review that combines several independent model runs in a production A/B test and reported improvements. That does not establish a benefit for every model combination: deployment choices should consider additional valid findings, false findings, verification effort, and cost.
Code review needs both coverage and useful comments
A pull request, or PR, proposes a change to a codebase. Developers inspect changed lines and the explanation of the change before deciding whether to merge it. A review agent participates at this stage by identifying defects and worthwhile improvements.
There are at least two questions about review quality: are the findings useful, and did the reviewer miss issues it should have found? Precision addresses the first; recall addresses the second.
Suppose a change has ten confirmed issues. An agent makes five findings, four of which correctly identify issues. Its finding precision is 4/5, or 80%, while its recall against the known issues is 4/10, or 40%. These are hypothetical numbers explaining the concepts, not scores for a product on ReviewBench.
Making more findings may reduce misses, but more false findings also force developers to inspect and dismiss them. ReviewBench lets readers separate usefulness from coverage instead of judging comment volume alone. In its actual scoring, the findings included in a calculation depend on how they match the reference set, as explained below.
219 PRs informed by a distribution of over 100 million
GitHub analyzed 103.9 million PRs to study distributions of languages, repository sizes, and change sizes. The benchmark itself contains 219 PRs across 187 public repositories and 19 languages, selected with that analysis in mind. Agents are not evaluated on all 103.9 million PRs.
Language and repository-size distributions track GitHub overall, while PR size is deliberately adjusted. The corpus favors the reviewable middle and tail so that tiny, single-file changes do not dominate. It is informed by real workloads, but it is not an unadjusted random sample of all GitHub activity.
Each evaluated change is fixed by base and head commits. Follow-up fixes can provide evidence of what was wrong, but the reviewed diff does not simply include those later fixes. Reviewing an already repaired change would not measure whether an agent could identify the original defect.
Official corpus construction and PR selection
Building the reference set and scoring a review
Before scoring agents, the benchmark builds a reference collection of findings, called the golden set. Despite the name, it is not a list of valid findings alone. It stores labels for useful findings, or true positives (TPs), and for incorrect or unhelpful findings, or false positives (FPs).
Candidate findings come from human reviews, inferred issues behind follow-up commits, deterministic tools such as static analyzers, and frontier models from multiple families. Findings describing the same underlying issue are semantically deduplicated. Agreement among several producers should not multiply the count of a single issue.
Collecting a candidate and accepting it as valid are separate operations. The current classifier is Claude Sonnet 5. Under a published rubric, it checks whether a finding is true, relevant, and non-trivial: specific enough to support a useful action. A human source or agreement among several models does not automatically make a finding correct.
REVIEWBENCH
Match AI findings against a shared golden set
Build the reference, then evaluate reviews of the same PRs.
Preparation
Build the reference
Collect candidates
Gather findings from people and tools.
Deduplicate and label
Merge the same underlying issues; use Sonnet 5 to label TP/FP, severity, and category.
Audit with human reviewers
Audit and correct the labels to finalize the golden set.
Golden set
A shared matching reference containing known TP and FP labels.
To matching in step 05For each evaluation
Evaluate the candidate agent
Review the fixed PR
Give the candidate the same code snapshot and collect its findings.
Match against the golden set
Check for the same underlying issue; classify unmatched findings with the grader.
Read the metrics
Inspect precision and recall, including severity and category breakdowns.
Grounded
Compare using existing golden-set labels.
Augmented
Also classify and score unmatched findings.
The methodology records that the corpus was initially labeled with Sonnet 4.6 and the classifier later moved to Sonnet 5. Human auditing included correcting 47 findings incorrectly labeled TP. The labeling stage in the diagram therefore does not mean accepting model output as final ground truth without review.
Matching is not exact string comparison. A model judges whether findings in the same file concern the same underlying issue. Different wording or somewhat different line ranges can still match, but the matcher itself can make mistakes.
Labeling, human auditing, and matching
Four metrics calculated from the same review
Even a broad golden set can omit issues nobody has found yet. ReviewBench separates grounded metrics, which use existing labels, from augmented metrics, which also classify unmatched findings. The precision denominator is particularly important: it differs between the two families.
| Metric | What it counts |
|---|---|
| Grounded precision | The fraction of matched candidate findings that inherit a TP label. Unmatched findings are excluded from both numerator and denominator. |
| Grounded recall | The fraction of golden TPs covered by the agent. Agents can be compared against the same known issues. |
| Augmented precision | Matched TPs plus unmatched findings newly classified TP, divided by all candidate findings. |
| Augmented recall | Covered golden TPs plus newly classified TPs, divided by golden TPs plus those new TPs. Each agent can have a different denominator. |
For a concrete example, suppose the golden set contains ten TPs as well as labeled FPs. An agent produces eight findings: four match known TPs and one matches a known FP. Three are unmatched; further classification labels two TP and one FP. This is a simplified hypothetical example without duplicates or findings that combine multiple issues.
| Scoring the same eight findings | Hypothetical calculation |
|---|---|
| Grounded precision | 4/5 = 80%. The three unmatched findings are excluded. |
| Grounded recall | 4/10 = 40%. Four of ten known TPs are covered. |
| Augmented precision | 6/8 = 75%. The two newly classified TPs also count. |
| Augmented recall | 6/12 = 50%. The denominator also gains the two new TPs. |
Here, augmented precision is lower than grounded precision even though the agent finds new valid issues. More findings enter the calculation, including an additional FP. Grounded precision alone therefore does not justify saying that 80% of all the agent’s comments are useful.
F1 combines precision and recall as twice their product divided by their sum. In this hypothetical example, grounded F1 is about 53.3% and augmented F1 is 60%. Each uses precision and recall from the same metric family. Fβ changes the preference: β below one emphasizes precision, while β above one emphasizes recall.
Grounded recall is the common reference for comparing coverage across agents. Augmented recall adds each agent’s newly classified TPs to the denominator, so different agents are assessed against different sets. It is useful as a diagnostic of discovery beyond the corpus, but ranking agents on it alone changes the basis of comparison.
Aggregation also distinguishes macro averages, which average per-PR metrics, from micro averages, which pool finding counts before calculating ratios. Macro treats PRs equally; micro is influenced more by PRs with more findings. Official evaluations average three repeated runs, so a single-PR example should not be confused with a leaderboard score.
Metric formulas, aggregation, and repeated runs
Severity and category explain what an overall score hides
Valid findings differ in impact: some prevent serious failures, while others make code easier to maintain. ReviewBench records severity separately from category so that readers can inspect the kinds of findings relevant to their needs.
The checked classifier rubric uses three severity levels: high, medium, and low. The announcement blog also uses “critical” for the top level; this article follows the scoring rubric’s “high” terminology. Categories include correctness, security, reliability, maintainability, and testing.
| Hypothetical illustration | What to examine |
|---|---|
| A change lets user A read user B’s private data | A security concern about a missing authorization check; potentially high impact if the exposure is substantial. |
| A change retries indefinitely after a network failure | A reliability concern about failure handling. Severity depends on the conditions and consequences. |
| A new name conflicts with established local conventions | A maintainability concern if there is a concrete local inconsistency, distinct from a purely stylistic preference. |
These examples were created by LATENT and do not reproduce corpus PRs or official labels. A security category does not automatically imply high severity. The affected users, triggering conditions, and existing protections all matter.
An agent designed to suppress low-severity findings may appear weak on overall recall. If its recall on high-impact issues is strong, it may suit a team that wants fewer notifications. A team seeking smaller improvements may value a different profile. Defining the use case first makes the breakdown more informative than an overall score alone.
Severity, categories, and TP/FP criteria
96.6% is judgment agreement, not a bug detection rate
GitHub says senior engineers who had not participated in dataset construction independently relabeled the findings. The reported 96.6% is agreement between human and benchmark TP/FP judgments. It does not mean that an agent found 96.6% of real bugs.
This is an internal audit of the benchmark organized by GitHub, not independent third-party certification of a product’s performance. Reviewers being independent of dataset construction is different from an evaluation being operated by an independent institution.
Agreement also depends on the label being judged. The methodology reports exact severity agreement of 62.9%. Strong agreement on TP versus FP does not mean equally strong agreement on how severe an issue is. Fine-grained comparisons by severity therefore also involve uncertainty in the labels.
Human–model agreement can still hide a shared mistaken assumption, and disagreement does not always mean the model is wrong. The official audit revisited the code to investigate which side was mistaken.
Audit results and classifier history
Combine model perspectives and measure added value and effort
One practical response to issues a single agent leaves unresolved is to use several models. The aim is to examine code from multiple perspectives and look for problems a single review missed. ReviewBench can help evaluate whether such a combination finds additional issues and whether it also adds false findings.
GitHub has evaluated this kind of combined review. Its lite-tier experiment for Copilot code review introduced a multi-model ensemble review that combines several independent model runs into a single review. GitHub reports that ReviewBench predicted increased precision, recall, and comment volume, together with lower cost per review, and that the production A/B test moved in the same directions.
| Production measure | Relative change versus production control |
|---|---|
| Addressed rate: comments judged to have led to a change | 8.0% increase |
| Recall measured through additional human review needed | 13.6% increase |
| Comment volume | 61% increase |
| Cost per review | 8.0% decrease |
The table reports GitHub’s production A/B results. Every value is a relative change versus the production control. An 8.0% increase does not mean an increase of 8.0 percentage points. The section does not provide the baseline absolute rates, so the resulting rates cannot be calculated from it.
For addressed rate, an LLM uses the diff, discussion thread, reactions, resolution state, and post-review code to judge whether a comment prompted a developer to change the code. It is a proxy for action taken, not a direct measure of finding correctness. Production recall is assessed through how much additional human review is needed; its calculation should not be assumed to be identical to ReviewBench recall with known TPs as the denominator.
GitHub’s reported production ensemble experiment
The result is evidence that the ensemble GitHub tested improved practical indicators. The section does not identify the model names and combination, number of model runs, detailed sample size, or experimental period. It does not establish that adding arbitrary different models will produce the same gains or lower costs. A 61% increase in comments also does not mean every added comment identified a valid new issue.
For a team’s own trial, one possible design is to let each reviewer work before showing it the others’ answers, then consolidate the findings. This is LATENT’s proposal, not a reconstruction of GitHub’s implementation. Give each model the same code revision and requirements so it can investigate without starting from another model’s conclusions.
Reviewers can also have different emphases: access control and input handling for security; empty inputs and limits for boundary conditions; and preservation of existing behavior and external contracts for regressions against requirements. Each reviewer uses the same code and requirements. A human should check the allocation of perspectives for gaps rather than relying on role names alone.
Consolidate findings by underlying cause and do not use a majority vote as the sole test of correctness. Three models can share the same mistaken assumption. A finding raised by only one model may still deserve verification against the code, requirements, reproduction conditions, and tests. Keeping the evidence for each finding makes it possible to distinguish duplicates from additional discoveries.
Decide whether to adopt the combination by comparing the final combined review with a single-agent review on the same PRs. Record additional known serious issues found, human time spent checking false findings and consolidating duplicates, latency, and execution cost. Some extra work may be worthwhile if it reduces serious misses; an increase in minor findings and checking effort alone is a reason to reconsider the models or their assignments.
Use augmented metrics to examine findings beyond the golden set, and grounded recall to compare coverage of known issues. Multiple perspectives are not an end in themselves: assess how much they reduce misses together with the verification work they add.
Deciding what to delegate in a real workflow
A useful starting point is to decide which burden the team wants to reduce: missed serious defects, or time spent checking false findings. For a payment-processing change, for example, a team could inspect recall on high-impact and correctness findings together with the corresponding precision. This is LATENT’s practical suggestion, not a reported benchmark result.
Next, try a small set of your own historical PRs. Human reviewers should examine both the agent’s findings and issues it did not mention. Checking only emitted findings reveals precision more readily than misses. Overlap with existing tests, waiting time, and cost also matter when deciding whether the review is worthwhile.
The official workflow offers a 25-PR test subset and three runs over all 219 PRs. The 25 PRs belong to the full corpus; they are not a separate hidden test. As tuning proceeds, performance on those examples should be distinguished from usefulness on other code.
GitHub’s participation workflow
What a public benchmark cannot settle
Public PRs make the evaluation inspectable, but 219 examples cannot establish equal performance on private business rules, non-public usage patterns, or organization-specific architectures. Broad language coverage still leaves a finite number of cases for each language and domain.
The corpus is not established as exclusively AI-authored code. GitHub has the ensemble evaluation and production experiment described above, but those are not a comprehensive controlled comparison of shared blind spots across writer-model and reviewer-model pairings. Whether a team’s own combination improves its workflow should be checked on its code and operating conditions.
The golden set also remains incomplete. Problems missed by every producer remain invisible. Augmented evaluation can recognize new findings, but their acceptance depends on the classifier. Automated grading cannot be assumed to fully repair shared blind spots.
Comparisons should hold the dataset, classifier, matcher, candidate version and configuration, and aggregation method constant. A change in grading can change a score even when findings stay the same. After repeated tuning on public examples, it is useful to check whether the pattern holds on fresh PRs. This is LATENT’s caution about over-adapting to a public test, not a finding that the corpus has contaminated any model’s training data.
Whether AI can safely inspect AI-written code depends on what is delegated. ReviewBench shows what an agent finds and what findings it produces under specified conditions. Combined with checks on your own code, that evidence can help define a useful supporting role. A high score alone does not establish the safety of merging or shipping code without human oversight.
Known limitations and reproducibility requirements
Sources and checked versions
Public materials were checked on October 7, 2026, Japan time. Reported benchmark and audit figures come from GitHub, not from LATENT running or independently scoring products. The hypothetical calculations and diagrams were created to explain the method.
- GitHub announcement: corpus scale, collection policy, and participation.
- Official methodology: formulas, auditing, matching, aggregation, and limitations.
- Official classifier prompt: TP/FP, severity, category, and other labeling criteria.
- ReviewBench website: entry point for public results and data.
- Official manifest of evaluated PRs.
The calculation method is pinned to the linked GitHub revision. The website and repository can reflect different update times, so detailed formulas follow that version of the methodology and classifier rubric. This article does not rank products because leaderboard results change over time and with evaluation conditions.