Abstract
Large language models have achieved remarkable capabilities across diverse domains, yet their safety alignment remains vulnerable to jailbreak attacks. In this work, we identify a previously underexplored failure mode - safety generalization lag - where alignment trained predominantly on natural language fails to transfer to the code domain. We show that this lag induces a code-completion blind spot, allowing malicious intent embedded within syntactically valid code to evade safety mechanisms. To exploit this vulnerability, we propose CodeMimicry, a fully automated black-box jailbreak framework that generates structured, object-oriented code prompts to induce harmful outputs via code completion. Experiments on 8 state-of-the-art commercial LLMs demonstrate that CodeMimicry achieves a 96.25% attack success rate with 1.51 queries on average, significantly outperforming both template-based and optimization-based baselines. Beyond empirical performance, we provide a mechanistic analysis of code-based jailbreaks through latent space representations, including projection onto refusal-related directions and activation steering. This analysis offers an explanation of how CodeMimicry bypasses safety mechanisms in code-related domains. Our findings reveal a weakness in current safety alignment and highlight the need for robust alignments in structured domains such as code.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Liang, Z., Huang, H., & Chen, W. (2026). CodeMimicry: Exploiting Safety Generalization Lag in Large Language Models via Structured Code Completion. https://omanscience.com/en/articles/codemimicry-exploiting-safety-generalization-lag-in-large-language-models-via-structured-code-completion
MLA 9
Liang, Zhen, et al. "CodeMimicry: Exploiting Safety Generalization Lag in Large Language Models via Structured Code Completion." https://omanscience.com/en/articles/codemimicry-exploiting-safety-generalization-lag-in-large-language-models-via-structured-code-completion.
Chicago (author–date)
Liang, Zhen, Hai Huang, and Wentao Chen. 2026. "CodeMimicry: Exploiting Safety Generalization Lag in Large Language Models via Structured Code Completion." https://omanscience.com/en/articles/codemimicry-exploiting-safety-generalization-lag-in-large-language-models-via-structured-code-completion.
Harvard
Liang, Z., Huang, H. and Chen, W. (2026) 'CodeMimicry: Exploiting Safety Generalization Lag in Large Language Models via Structured Code Completion', Available at: https://omanscience.com/en/articles/codemimicry-exploiting-safety-generalization-lag-in-large-language-models-via-structured-code-completion.
Vancouver
Liang Z, Huang H, Chen W. CodeMimicry: Exploiting Safety Generalization Lag in Large Language Models via Structured Code Completion. https://omanscience.com/en/articles/codemimicry-exploiting-safety-generalization-lag-in-large-language-models-via-structured-code-completion
IEEE
Z. Liang, H. Huang, and W. Chen, "CodeMimicry: Exploiting Safety Generalization Lag in Large Language Models via Structured Code Completion," https://omanscience.com/en/articles/codemimicry-exploiting-safety-generalization-lag-in-large-language-models-via-structured-code-completion.