Controlled Decoding Attacks on Black-Box LLMs

Researchers propose a method for jailbreaking large language models (LLMs) through text-only interfaces, using a framework that reconstructs probabilities from sampled outputs and selectively modifies the distribution. This affects the safety alignment of LLMs and has implications for their security and reliability.

RSS Score 0 9/30/2026, 4:00:00 AM Original Source
Save an API key to vote.