Hard-Gate Candidacy in a Deployed Validator Suite

This paper describes an evaluation of a deployed validator suite, specifically the hard-gate candidacy of 13 validators in a generative agent. The study tests the validators against 900 builds labeled by downstream outcome and reports the marginal separation of each check. The results show that some checks are not distinguishable from zero and that a skipped check is recorded as a pass, imposing a ceiling on check quality.

RSS Score 0 10/1/2026, 4:00:00 AM Original Source
Save an API key to vote.