ShieldCLIP: Selective Safety Alignment for Harmful Content Mitigation in Multimodal Foundation Models

This paper introduces ShieldCLIP, a framework for selective safety alignment in multimodal foundation models like CLIP. It conditions safety alignment on the observed safety state of each modality, preserving safe content and redirecting only unsafe content. The authors also introduce ViSUv2, a 195k-quadruplet dataset with independent per-modality safety labels. ShieldCLIP achieves consistent reductions in harmful outputs in various tasks and settings.

RSS Score 0 10/1/2026, 4:00:00 AM Original Source
Save an API key to vote.