CounterSteer: Suppressing Indirect Prompt Injection with Activation Steering
CounterSteer is a defense against indirect prompt injection in LLMs, suppressing the behavior by subtracting a learned direction from tool-result tokens during prefill. It requires no fine-tuning, auxiliary models, or added tokens, and achieves 93-100% typography-normalized benign utility while reducing attack success rates to 0.00-0.17.
Save an API key to vote.