Abstract
Emotional expression can influence the safety decisions of large language models (LLMs), offering a potential avenue for improving safety alignment. Existing studies have mainly focused on how emotional expressions facilitate attacks under harmful requests, while overlooking their effects on benign requests. We find that emotional expression can also systematically increase refusal tendencies on benign requests, leading to unnecessary over-refusal. Based on this observation, we propose emotion-guided refusal subspace steering (EmoRSS), an activation-steering method that mitigates emotion-induced over-refusal while preserving refusal behaviour on harmful requests. Specifically, we first identify a refusal-sensitive layer using layer-wise linear probes and construct a refusal subspace from sparse autoencoder (SAE) features aligned with the probe direction. Next, we use paired regular and emotional requests with the same queries to estimate the mean activation shift in the features defining the refusal subspace. Finally, we decode this shift into an activation intervention vector and apply it in the reverse refusal direction during inference, without updating the backbone parameters. Experiments on two LLMs show that, when requests contain emotional expressions, our method achieves a more favourable trade-off between refusing harmful requests and answering benign ones than prior over-refusal mitigation baselines, while better preserving general task performance.
Keywords
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Miao, S., Ma, Y., Cui, C., Liu, X., Jisheng, D., Zhuo, S., Shen, F., & Chua, T. S. (2026). EmoRSS: Mitigating Emotion-Induced Over-Refusal in Large Language Models. https://omanscience.com/en/articles/emorss-mitigating-emotion-induced-over-refusal-in-large-language-models
MLA 9
Miao, Shuyi, et al. "EmoRSS: Mitigating Emotion-Induced Over-Refusal in Large Language Models." https://omanscience.com/en/articles/emorss-mitigating-emotion-induced-over-refusal-in-large-language-models.
Chicago (author–date)
Miao, Shuyi, Yaojin Ma, Chenhang Cui, Xiaohao Liu, Dang Jisheng, Shengda Zhuo, Fei Shen, and Tat-Seng Chua. 2026. "EmoRSS: Mitigating Emotion-Induced Over-Refusal in Large Language Models." https://omanscience.com/en/articles/emorss-mitigating-emotion-induced-over-refusal-in-large-language-models.
Harvard
Miao, S., Ma, Y., Cui, C., Liu, X., Jisheng, D., Zhuo, S., Shen, F. and Chua, T. S. (2026) 'EmoRSS: Mitigating Emotion-Induced Over-Refusal in Large Language Models', Available at: https://omanscience.com/en/articles/emorss-mitigating-emotion-induced-over-refusal-in-large-language-models.
Vancouver
Miao S, Ma Y, Cui C, Liu X, Jisheng D, Zhuo S, et al. EmoRSS: Mitigating Emotion-Induced Over-Refusal in Large Language Models. https://omanscience.com/en/articles/emorss-mitigating-emotion-induced-over-refusal-in-large-language-models
IEEE
S. Miao, Y. Ma, C. Cui, X. Liu, D. Jisheng, S. Zhuo, F. Shen, and T. S. Chua, "EmoRSS: Mitigating Emotion-Induced Over-Refusal in Large Language Models," https://omanscience.com/en/articles/emorss-mitigating-emotion-induced-over-refusal-in-large-language-models.