ThinkingGuard: Decoding Implicit Hazards via Step-by-Step Risk Attribution in Multimodal Large Language Models

ThinkingGuard is a guard model for identifying implicit hazards in Multimodal Large Language Models (MLLMs). It uses a step-supervised structured reasoning framework and a step-reward Monte Carlo Tree Search algorithm to detect risks in MLLMs. The authors propose a new dataset, TriggerBench, to train and evaluate ThinkingGuard. The system demonstrates strong performance on both standard and implicit safety benchmarks.

RSS Score 0 9/30/2026, 4:00:00 AM Original Source
Save an API key to vote.