UniGuardian: A Unified Defense for Detecting Prompt Injection, Backdoor Attacks and Adversarial Attacks in Large Language Models

Researchers propose UniGuardian, a training-free detector for Large Language Models (LLMs) that identifies prompt injection, backdoor, and adversarial attacks without knowing the attack type. UniGuardian measures how prompt perturbations shift the model's output distribution and uses a single-forward strategy for efficient detection and text generation.

RSS Score 0 10/2/2026, 4:00:00 AM Original Source
Save an API key to vote.