GitHub Blog 🛠 Dev 👁 0 📖 9 min read

ReviewBench: An open benchmark for AI code review

We’re launching ReviewBench, a benchmark for code review agents built on representative GitHub pull requests, multi-source ground truth, calibrated evaluation, and production-aligned metrics. The post ReviewBench: An ope

Agentic code review is becoming an essential piece of how development happens. It helps you inspect pull requests, catch issues, and decide what deserves attention before code ships.

But the quality of existing AI reviewers can be hard to measure, and you need to know the strengths of a reviewer before you know if it will help you. Some reviewers surface more issues, some produce less noise, and some are stronger at catching critical problems while others surface smaller improvements, too. You may need code review to do different things within your workflow.

That makes it important to understand how reviewers actually compare: what different systems catch, what they miss, and the tradeoffs they make. A good code review benchmark should reflect the diversity of real pull requests, capture a broad set of review findings, and support meaningful breakdowns by severity, category, and precision-recall preferences. For teams building code review agents, the benchmark should also provide an offline signal that reliably tracks whether changes are likely to improve the experience in production. Existing benchmarks often make tradeoffs between label quality, coverage, and how well they represent real-world code review, leaving a gap for a rigorous and reproducible evaluation methodology that brings these pieces together.

We built ReviewBench, a new code review offline benchmark, to address that gap, and it is available for you to use today. It follows the language, repo size, and size distribution of pull requests, modeled after over 100 million real pull requests on GitHub. It uses a multi-source golden set and a consistent evaluation rubric and has been independently validated by senior engineers. Just as important, with the help of ReviewBench, our offline evaluation of Copilot code review (CCR) has become more effective at anticipating the direction of production experiments, giving us greater confidence that measured improvements reflect meaningful gains for users.

In this post, we’ll walk through how ReviewBench is constructed, how it establishes reliable ground truth and scoring, and how to onboard your own code review system and submit results.

Definitions of terms used in this blog post

  • Benchmark: A standardized evaluation that tests code reviewers on a common set of pull requests using the same scoring methodology.
  • Finding: A specific issue surfaced during code review.
  • Golden set: A validated collection of known findings for each pull request, used as a reference for evaluating what a reviewer catches or misses.
  • Precision: Of the issues a reviewer surfaces, the proportion that are valid. Higher precision generally means less noise.
  • Recall: Of the known valid issues, the proportion the reviewer finds. Higher recall means broader coverage.
  • F1 score: A single score that balances precision and recall equally.
  • Fβ score: A variation of F1 that lets you put more weight on either precision or recall, depending on your review preference.

ReviewBench at a glance

1

What we built

A realistic, comprehensive benchmark for AI code review agents

103.9M

GitHub pull requests

Analyze distributions by language, repository size, and change shape.

Representative benchmark corpus

219 public pull requests across 19 languages, aligned to GitHub-wide distributions while preserving substantive review cases.

Multi-source golden set

  • Human reviewers
  • Frontier LLMs
  • Static analysis

Structured findings

Every finding is labeled for severity and category, enabling user-tailored slices.

Severity

  • Critical
  • Medium
  • Low

Category

  • Correctness
  • Security
  • Reliability
  • Maintainability
  • Testing
  • ......

Evaluation metrics

Four metrics measure both known and newly discovered issues.

  • Grounded precision
  • Grounded recall
  • Augmented precision
  • Augmented recall

Objective evaluation

Measure improvement and compare across agents objectively. Help users choose the reviewer that fits their needs the best.

2

How we keep it trustworthy

An auditable chain from rubric to expert validation and production checks

Published rubric

One explicit standard for all findings.

Human-labeled dev set

Senior engineers establish ground truth.

Calibrated grader

Aligned with human judgment.

Uniform labeling

Same standard across all sources.

Published agreement

Expert audit of benchmark quality.

Auditable end to end

96.6% agreement

Senior engineers independently labeled golden true-positives before release.

Offline signals that anticipate production

Benchmark movement is checked against online experiments.

  • Improvements tend to show up online
  • Regressions tend to show up online too

How ReviewBench works

Our benchmark is built around five principles:

1. Representative pull requests, not a demo set

We analyzed 103.9 million GitHub pull requests to characterize the real-world distribution of code review workloads. ReviewBench contains 219 pull requests from 187 public open source licensed repositories spanning 19 languages, with its language and repository-size distributions closely matching GitHub overall. The complete benchmark dataset is publicly available.

We make one deliberate adjustment to this distribution: while language and repository size mirror GitHub directly, pull request size is weighted toward the reviewable middle and tail. This reduces the overrepresentation of tiny, single-file changes while preserving more substantive, multi-file pull requests where review quality matters most.

Quick corpus snapshot:

2. Broad ground truth discovery, independently judged

No single reviewer, whether human or model, can identify everything worth finding in a pull request. To build a broader and more reliable golden set for ground truth findings, we follow a three-stage process:

  • Gather candidate findings from diverse sources. We collect findings from real human reviewers, issues inferred from author follow-up commits, deterministic analysis tools, and multiple frontier LLMs across model families.
  • Semantically deduplicate overlapping findings. We merge findings that identify the same underlying issue, broadening coverage without allowing agreement across producers to artificially inflate the golden set or making it dependent on any one source’s blind spots.
  • Validate findings under a shared rubric. The source of a finding does not determine whether it is correct: a finding counts as a true positive only if it is true, relevant, and non-trivial. We use Claude Sonnet 5 as the LLM grader, applying a consistent evaluation rubric across all submissions. For transparency and reproducibility, we publish both the evaluation rubric and the judge used to apply it.

3. Metrics that measure both known and newly discovered issues

Most benchmarks report precision and recall against a fixed golden set. ReviewBench reports six metrics in two families:

  • Grounded precision, recall, and F1 score use only the existing gold-set labels. They provide the strict, apples-to-apples comparison: of the issues we already know about, how many did the agent find, and what share of its findings matched a known issue?
  • Augmented precision, recall and F1 score also evaluate findings that do not match anything in the golden set. The judge independently determines whether those unmatched findings are true or false positives, allowing a reviewer to receive credit for valid issues that no producer in the golden set surfaced

That distinction becomes more important as review agents become more capable. A fixed golden set inevitably becomes incomplete as systems discover issues its creators did not anticipate. Augmented metrics let ReviewBench recognize that behavior rather than automatically penalizing it. Because augmented recall expands the denominator based on what each agent discovers, we use grounded recall as the headline cross-system comparison and augmented metrics as an additional per-system diagnostic.

4. Configurable evaluation for different review preferences

There is no single universally optimal review experience. Some developers may want to focus only on critical issues, while others also value lower-severity, non-breaking findings. Some prefer broader coverage, while others prioritize precision and minimal noise. Others may have specialized needs, such as security- or privacy-focused review.

ReviewBench lets results be sliced by severity and category, while precision and recall capture different operating preferences. Users can also adjust β in the Fβ score to place more weight on recall for broader coverage or precision for lower noise. As these preferences change, the leaderboard is re-ranked accordingly, helping users identify the systems that best match their review priorities.

5. Internally audited and reproducibly evaluated

Before release, we asked senior engineers who had not participated in building the benchmark dataset to independently re-label every ground-truth finding from scratch. Their true/false-positive judgments agreed with ReviewBench 96.6% of the time. We version the benchmark dataset, judge, and matcher used in every evaluation, so results can be compared under the same benchmark configuration and revalidated when the benchmark changes. We also publish the validation methodology, agreement measurements, and known threats to validity, so readers can see how benchmark quality is assessed and where uncertainty remains.

Explore ReviewBench

ReviewBench’s research preview version is now available through the ReviewBench website, where you can explore the full benchmark, compare code review agents, and bring your own agent to evaluate and iterate.

With ReviewBench, you can:

  • Explore the full benchmark dataset. The complete ReviewBench dataset is publicly available, including the pull requests, findings, labels, severity and category annotations. This allows you to inspect exactly what systems are evaluated on and reproduce benchmark results.
  • Compare systems on the leaderboard. Results from evaluated code review agents using the full benchmark data are published on a common leaderboard, with views across overall performance, severity, category, and different precision–recall preferences.
  • Bring your own agent and hill-climb. The full benchmark dataset, evaluation methodology, LLM judge prompt, judge model configuration, and self-serve runner are publicly available, so you can evaluate your own code review agent, inspect its strengths and gaps, and iterate against the same benchmark configuration.

How we’ve used ReviewBench

We have used ReviewBench to evaluate Copilot code review (CCR) across successive iterations, giving us a consistent way to measure progress, catch regressions, and prioritize promising changes. Over time, this has helped us improve the product. One of the most valuable benefits of ReviewBench is that it provides an early offline signal of how a change to the product is likely to perform in production. Across experiments evaluated with ReviewBench before A/B testing, offline changes have consistently pointed in the same direction as what we see later in production.

A recent lite-tier experiment provides a concrete example of this broader pattern. We introduced a multi-model ensemble review that combines several independent model runs into a single review rather than relying on a single run. ReviewBench predicted higher precision, recall, and comment volume, along with lower cost per review.

To compare offline and production results, we use corresponding online signals. Addressed rate, our online counterpart to precision, is the percentage of CCR comments that an LLM determines prompted a developer to make a corresponding code change, based on the diff, thread, reactions, resolution state, and post-review code. For recall, we measure how much additional human review is still needed.

The online A/B test moved in the same direction as ReviewBench predicted: addressed rate (precision) rose 8.0%, recall rose 13.6%, and comment volume rose 61%, while cost per review fell 8.0%, all relative to the production control.

Comment volume alone, however, does not capture comment quality. More critical findings mean something very different from low-severity nits. ReviewBench’s severity-level evaluation captured this too: it predicted a 227% increase in critical comments, compared with 262% online, along with the same broader shift toward more moderate comments and fewer nits.

This gives us a fast and repeatable signal before running production experiments. Online experiments remain the ultimate measure of user impact, but ReviewBench gives us greater confidence in which changes are worth taking there.

How to submit your own run

  1. Sign in with GitHub on the ReviewBench website.
  2. Register your agent. Provide a container image, your configuration, and your own model key. We provide the judge.
  3. Try it on the test set. Run against a 25-PR test set with per-PR detail and repeat as you tune your configuration.
  4. Do a final run. When you’re ready, run the full set of 219 pull requests (three rounds), scored by the same judge as every other entry.
  5. Publish to the leaderboard. Your scores remain private until a maintainer reviews and approves the submission. Scores are published to the leaderboard only if they outperform the agent’s current leaderboard score, or if this is the agent’s first leaderboard entry.

We invite you to explore ReviewBench, evaluate your own system, challenge our assumptions, and help us improve the benchmark. We’re excited to collaborate with researchers and practitioners to make code review evaluation more open, reliable, and useful—and ultimately help move AI code review forward.

Acknowledgments

ReviewBench was a team effort across GitHub and Microsoft. We’re grateful to the researchers and engineers who built it: those who designed the methodology, curated the pull requests, built the golden set and the evaluation pipeline, and made the benchmark something anyone can run.

The post ReviewBench: An open benchmark for AI code review appeared first on The GitHub Blog.

📰 Read the original article on GitHub Blog

Originally published by GitHub Blog. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.