Same Bytes, Different Authority: Reserved-Token Representations in Chat-Template Prompt Injection
Researchers found a way to strengthen prompt injection attacks against LLM agents by wrapping injected instructions in the model's own chat template. This allows attackers to evade tokenization-based defenses. The study measured the effectiveness of this technique on various LLM models and found significant improvements in attack success rates.
Save an API key to vote.