الباحثون

Yunpeng Zhai

المنشورات 1

نسخة أولية وصول مفتوح

Playing social deduction games with reinforcement fine-tuned large language models

Lingzhe Zhang, Yunpeng Zhai, Tong Jia وآخرون · 2026

Reinforcement fine-tuning (RFT) is increasingly used in applications where large language models (LLMs) interact with humans and other agents. Here we use social deduction games to study how RFT changes LLMs' social behaviour. We let fine-tuned and base LLM agents play hidden-role games that require hidden-state infere …

المؤلفون المشاركون