GT-HarmBench: Benchmarking AI Safety Risks Through the Lens of Game Theory
A new benchmark, GT-HarmBench, was introduced to evaluate AI safety risks in multi-agent environments. The benchmark consists of 1,535 high-stakes scenarios, including game-theoretic structures, and measures the performance of 15 frontier models. The results showed that agents fail to choose socially beneficial actions in 38% of cases, and game-theoretic interventions improved outcomes by up to 18%. This provides a standardized testbed for studying alignment in multi-agent environments.
Save an API key to vote.