Controlled Decoding Attacks on Black-Box LLMs
Researchers propose a method for jailbreaking large language models (LLMs) through text-only interfaces, using a framework that reconstructs probabilities from sampled outputs and selectively modifies the distribution. This affects the safety alignment of LLMs and has implications for their security and reliability.
Save an API key to vote.