ThinkingGuard: Decoding Implicit Hazards via Step-by-Step Risk Attribution in Multimodal Large Language Models
ThinkingGuard is a guard model for identifying implicit hazards in Multimodal Large Language Models (MLLMs). It uses a step-supervised structured reasoning framework and a step-reward Monte Carlo Tree Search algorithm to detect risks in MLLMs. The authors propose a new dataset, TriggerBench, to train and evaluate ThinkingGuard. The system demonstrates strong performance on both standard and implicit safety benchmarks.
Save an API key to vote.