نسخة أولية وصول مفتوح
The Model Plants the Trigger: Answer-Side Backdoor Attacks in Multi-Turn Large Language Models
Safety alignment in Large Language Models (LLMs) remains vulnerable to backdoor attacks. Existing LLM backdoors are almost all input-centric: activation depends on explicit trigger patterns in the user input, so modern guardrails are built to sanitize the input space. We challenge this assumption with a novel answer-si …