CounterSteer: Suppressing Indirect Prompt Injection with Activation Steering

CounterSteer is a defense against indirect prompt injection in LLMs, suppressing the behavior by subtracting a learned direction from tool-result tokens during prefill. It requires no fine-tuning, auxiliary models, or added tokens, and achieves 93-100% typography-normalized benign utility while reducing attack success rates to 0.00-0.17.

RSS Score 0 9/30/2026, 4:00:00 AM Original Source
Save an API key to vote.