Backdoor Purification for LoRA-Tuned LLMs via Null-Space Projection
Researchers propose a method to purify LoRA-tuned LLMs from backdoor attacks without prior knowledge of triggers or access to clean references, reducing attack success rates from nearly 100% to less than 10%.
Save an API key to vote.