GitHub announces ReviewBench for AI code-review agents
GitHub announced ReviewBench, an open benchmark for evaluating AI code-review agents on real-world pull requests. The benchmark is available in research preview through its website.

Key takeaways · 3
- 01
Teams can explore ReviewBench in research preview and inspect its public corpus manifest and golden findings.
- 02
The benchmark’s main comparison measure is grounded recall; users can also filter by severity and category and adjust β in the Fβ score.
- 03
Agent submissions include a 25-pull-request test set before three full-set evaluation rounds across 219 pull requests.
A benchmark built around real pull requests
GitHub announced ReviewBench as a benchmark for AI code-review agents, describing it as an open, reproducible way to evaluate systems on real-world pull requests.[1][2] The benchmark runs agents against the same 219 pull requests, drawn from 187 public repositories and spanning 19 programming languages.[3] Pull-request sizes are weighted toward the reviewable middle and tail.[3] GitHub’s announcement was authored by Michelle Zhou and Alejandro Carderera de Diego and published on October 5, 2026.[1] ReviewBench is available in research preview through its website.[3]
How the benchmark judges findings
ReviewBench is designed to measure whether an AI reviewer finds issues while avoiding false alarms.[3] GitHub built its reference set by collecting candidate findings from human reviewers, code changes, analysis tools and AI models, then merging and assessing those candidates.[3] The benchmark uses grounded recall as its main measure for comparing systems.[3] Its F1 score gives precision and recall equal weight, while users can filter results by severity and category and adjust β in the Fβ score.[3] An AI judge assesses whether additional findings are valid.[3]
Submission and access details
To register an agent, users sign in with GitHub and provide a container image, configuration and model access key.[3] The submission process offers a test set of 25 pull requests before three full-set evaluation rounds on the 219-pull-request benchmark.[3] Submitted scores remain private until a maintainer approves them.[3] GitHub says the benchmark continues to run if an upstream repository is deleted or rewritten.[2] The ReviewBench repository is released under the MIT license, and the full corpus manifest and golden findings are public.[2]
GitHub reports results from a Copilot test
GitHub uses ReviewBench to evaluate Copilot code-review changes before testing them with users.[3] In a production test, GitHub reported that the share of review comments judged by an AI model to have prompted code changes rose by 8%, while recall rose by 13.6%.[3] GitHub also reported that cost per review fell by 8% relative to the production control.[3] Feedback in the test shifted toward critical and moderate issues, with fewer minor suggestions.[3] These figures describe that production test; the evidence does not establish that the same results will apply to other agents or teams.[3]
ReviewBench is described as an open, reproducible benchmark for evaluating AI code-review systems on real-world pull requests. Its full corpus manifest and golden findings are public, and submitted scores remain private until a maintainer approves them.
Why it matters
Put this to work — one session a day, built for your industry.
Create a free account for a daily session — eight questions and one real-work challenge, on the news that affects your role.
Start freeHow this developed
7 October 2026
GitHub announces ReviewBench for AI code-review agents
Sources
- ReviewBench: An open benchmark for AI code review - The GitHub Bloggithub.blog
- GitHub - review-bench/ReviewBench: ReviewBench is an open, reproducible benchmark for evaluating AI code review systems on real-world pull requests. · GitHubgithub.com
- GitHub’s ReviewBench puts AI code reviewers to the test - Help Net Securityhelpnetsecurity.com