Current LLMs struggle with multi-round code review—their performance drops significantly as review iterations increase, they miss complex defects, and they fail to track how issues change across multiple rounds of feedback.
MCR-Bench is a new benchmark for evaluating AI models on realistic code review tasks that involve multiple rounds of back-and-forth interaction between developers and reviewers.