JevAdvBench: A Benchmark and Black-Box Attacks for Reinforcement Learning for Calibrated Decisions Models
A new benchmark (JevAdvBench) has been developed to measure the robustness of reinforcement learning for calibrated decisions (RLCD) models, specifically targeting models like Jev. The benchmark and black-box attacks can help identify vulnerabilities in these models, which are used in applications where software acts on model answers without human review.
Save an API key to vote.