Depth, Not Breadth: Best-of-N Jailbreaking Beyond Surface Noise
Researchers demonstrate a new attack method called best-of-N jailbreaking against self-check defenses, specifically targeting code-completion encodings used in LLMs. They show that by moving variance into a structural channel, the attack can reach 67% of behaviors, outperforming the strongest published self-check defense. This highlights the importance of separating screening and generation in LLM architecture.
Save an API key to vote.