Guard Models Are Overconfident Where Base Models Are Uncertain
Researchers evaluated five guard models for prompt classification and found that they can be overconfident when faced with adversarial attacks, leading to high-confidence errors. This mismatch between guard confidence and base model uncertainty can have significant implications for safety classifiers in AI systems.
Save an API key to vote.