UniGuardian: A Unified Defense for Detecting Prompt Injection, Backdoor Attacks and Adversarial Attacks in Large Language Models
Researchers propose UniGuardian, a training-free detector for Large Language Models (LLMs) that identifies prompt injection, backdoor, and adversarial attacks without knowing the attack type. UniGuardian measures how prompt perturbations shift the model's output distribution and uses a single-forward strategy for efficient detection and text generation.
Save an API key to vote.