Trending: On-device modelsSearch
iHeartGeek
iTECH

GitHub's ReviewBench puts AI code reviewers to the test

GitHub has published ReviewBench, an open benchmark built from 219 real pull requests that scores code review agents on precision and recall, with a published rubric and an audited golden set.

A translucent glass cube with a diamond lattice pattern glowing violet and magenta, lit from within and centred on a dark reflective surface against a deep blue and magenta background

GitHub has published ReviewBench, an open benchmark for the agents that now review code before it ships. The company announced it on 5 October, alongside the evaluation methodology, the scoring rubric and the dataset, and says the benchmark is available to use now.

Built from real pull requests, not a demo set

The corpus is deliberately unglamorous. GitHub analysed 103.9 million pull requests to characterise what review work actually looks like, then assembled a benchmark of 219 pull requests drawn from 187 public open-source repositories across 19 languages, with the language and repository-size mix matched to GitHub overall. One distribution is weighted on purpose: pull request size is skewed towards the reviewable middle and tail, so tiny single-file changes do not crowd out the multi-file work where review quality matters most.

A golden set with more than one author

No single reviewer, human or model, finds everything worth finding, so GitHub builds its ground truth from several sources: real human reviewers, issues inferred from follow-up commits by the original author, deterministic static analysis, and a spread of frontier models from different families. Overlapping findings are semantically deduplicated so that agreement between producers cannot inflate the set, and every candidate is then validated against a shared rubric. A finding counts only if it is true, relevant and non-trivial. Claude Sonnet 5 acts as the grader, and both the rubric and the judge are published so the result can be reproduced.

Six metrics, and credit for finding something new

Most code review benchmarks report precision and recall against a fixed answer key. ReviewBench splits that into two families. Grounded precision, recall and F1 score only count findings that match the golden set, which gives a strict like-for-like comparison between agents. Augmented precision, recall and F1 also consider findings that match nothing in the set, with the judge deciding whether each is genuinely valid, so a reviewer gets credit for real issues that no producer in the golden set surfaced — a distinction that matters as agents get better at spotting things their benchmark was never built to contain. Results can be sliced by severity and category, from critical correctness and security findings down to low-severity maintainability and testing notes, and the F-beta score can be tuned towards precision or recall.

Trust, and a signal that tracks production

GitHub's trust argument rests on auditability: a published rubric, a human-labelled development set built by senior engineers, a grader calibrated against human judgement, the same labelling standard applied to every source, and a published agreement figure. On that last point GitHub reports 96.6 per cent agreement, from senior engineers independently labelling golden true positives before release. The company also says ReviewBench has made its offline evaluation of Copilot code review better at predicting which changes will actually help users in production, which is the test that matters for anyone shipping a review agent. Because the benchmark is open, other teams can onboard their own reviewer and submit results for comparison.

Our opinion

The quiet admission buried in ReviewBench is that a fixed answer key is a losing proposition. If review agents keep improving, a golden set assembled today becomes a ceiling tomorrow, and a benchmark that only rewards matching known findings will eventually punish the systems that are best at finding the unknown ones. Splitting the score into grounded and augmented halves is GitHub conceding that problem in the open, and it is a more honest arrangement than most vendor benchmarks manage. The 96.6 per cent agreement figure deserves the scepticism any self-reported number earns, and the deeper question is whether a 219-pull-request sample can carry the weight of a public ranking. Still, a rubric you can read, a judge you can inspect and a dataset you can download is a meaningful improvement on the alternative, which is everyone trusting whatever chart their vendor put in a launch post.