MMSkillRisk: Can Agents Stay Safe When Multimodal Skills Become Traps?
Researchers introduced MMSkillRisk, a benchmark for evaluating image-borne attacks in multimodal skills. They designed Native-Context Visual Attack (NCVA), which disguises malicious instructions as native components of teaching images. The attack was successful in 43.1% of cases, with higher success rates in certain configurations. This highlights the risk of skill-bundled images inducing unauthorized actions.
Save an API key to vote.