CRJudgeBench: Can AI Detect Plausible but Invalid Code Reviews?

Researchers introduced CRJudgeBench, a benchmark for evaluating AI's ability to detect technically incorrect code-review comments. They also presented Sentinel, a repository-grounded agentic judge that uses iterative action-level learning to improve accuracy. The study showed that even state-of-the-art LLMs struggle to identify untrustworthy comments, but Sentinel outperformed its base model and a competitor model by a significant margin.

RSS Score 0 9/30/2026, 4:00:00 AM Original Source
Save an API key to vote.