الباحثون

Xianhui Zhang

المنشورات 1

نسخة أولية وصول مفتوح

ACTR: Aligning Thoughts and Responses for Multilingual Safety in Reasoning LLMs

Xianhui Zhang, Jian Yu, Chengyu Xie وآخرون · 2026

Ensuring the safety of reasoning large language models (LLMs) across languages is essential for their reliable deployment. However, when exposed to jailbreak attacks in non-high-resource languages, these models may generate unsafe responses even when their reasoning traces identify safety risks. To address this issue, …

المؤلفون المشاركون