CodeMimicry: Exploiting Safety Generalization Lag in Large Language Models via Structured Code Completion
A research paper proposes a framework called CodeMimicry, which exploits a safety generalization lag in large language models by generating structured code prompts to induce harmful outputs via code completion, achieving a 96.25% attack success rate on 8 state-of-the-art commercial LLMs. This highlights a weakness in current safety alignment and emphasizes the need for robust alignments in structured domains like code.
Save an API key to vote.