AI news · October 7, 2026
GitHub publishes ReviewBench for testing AI code review
ReviewBench turns AI code review into a public measurement problem. Its 219-PR preview gives builders a way to compare grounded findings instead of counting comments or accepting vendor scores.
- public pull requests
- 219
- programming languages
- 19
- senior-engineer agreement
- 96.6%
GitHub published ReviewBench on October 5 as an open benchmark for AI-assisted code review. The project was modelled on 103.9 million GitHub pull requests and currently uses 219 public pull requests from 187 repositories across 19 programming languages. Its reference set combines human review, large-language-model review, and static-analysis evidence so a finding can be checked against the changed code and the issue it claims to identify.
The preview reports grounded precision and recall, plus augmented precision and recall. GitHub says senior engineers agreed with the reference findings at a 96.6 percent rate. The company used the benchmark to evaluate Copilot code review and says offline changes pointed in the same direction as an online A/B test. The dataset is still small compared with real repository diversity.
The shift is from demo quality to reproducible review quality. A code-review product must show which findings are real, actionable, and missed. Teams still need a private holdout set for their languages and security rules.
What you can do with it
Run your review agent on the public 25-PR trial, then create a private holdout from closed defects and accepted review comments. Track grounded precision, missed defects, developer acceptance, and time saved. Treat comment volume as a failure metric when it does not improve merged code.
Our take
Open benchmarks are more valuable here than another polished code-review demo. ReviewBench is not proof that any tool is production-ready, but it gives buyers and builders a common starting point for testing false positives and missed defects before paying for scale.
Links GitHub announcement
Source: GitHub ↗ — Made With Models writes the brief; the reporting is theirs.