Abstract
Building LLMs that behave well socially, not merely correctly, requires Building LLMs that behave well socially, not merely correctly, requires more than producing locally helpful responses. A socially competent agent must infer users' unstated goals, respect their preferences, and adapt as the conversation unfolds. These behaviors are inherently multi-turn and social, making them hard to optimize: real interaction data is scarce, and user preferences are typically latent rather than directly observable. To address these challenges, we build on a persona-driven social simulation environment (consisting of a persona library, LLM-based user simulators, and a user-satisfaction scoring system ranging from [0, 1]), to introduce preference-batched GRPO (PB-GRPO), a post-training algorithm that learns socially adaptive policies from conversation-level feedback. Compared to vanilla GRPO, PB-GRPO computes advantages using a normalization estimated across a bucket of users with similar preferences, stabilizing training across a diverse social population. Empirical evidence shows that PB-GRPO improves models' social behavior over strong reinforcement learning baselines in our simulated environment.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Wang, J., Yin, J., Han, X., Mei, Y., Hao, J., & Guo, B. (2026). PB-GRPO: Learning Socially Adaptive LLM Agents from Persona-Driven Simulation with Preference-Batched GRPO. https://omanscience.com/en/articles/pb-grpo-learning-socially-adaptive-llm-agents-from-persona-driven-simulation-with-preference-batched-grpo
MLA 9
Wang, Jingquan, et al. "PB-GRPO: Learning Socially Adaptive LLM Agents from Persona-Driven Simulation with Preference-Batched GRPO." https://omanscience.com/en/articles/pb-grpo-learning-socially-adaptive-llm-agents-from-persona-driven-simulation-with-preference-batched-grpo.
Chicago (author–date)
Wang, Jingquan, Jun Yin, Xu Han, Yongsheng Mei, Jie Hao, and Bin Guo. 2026. "PB-GRPO: Learning Socially Adaptive LLM Agents from Persona-Driven Simulation with Preference-Batched GRPO." https://omanscience.com/en/articles/pb-grpo-learning-socially-adaptive-llm-agents-from-persona-driven-simulation-with-preference-batched-grpo.
Harvard
Wang, J., Yin, J., Han, X., Mei, Y., Hao, J. and Guo, B. (2026) 'PB-GRPO: Learning Socially Adaptive LLM Agents from Persona-Driven Simulation with Preference-Batched GRPO', Available at: https://omanscience.com/en/articles/pb-grpo-learning-socially-adaptive-llm-agents-from-persona-driven-simulation-with-preference-batched-grpo.
Vancouver
Wang J, Yin J, Han X, Mei Y, Hao J, Guo B. PB-GRPO: Learning Socially Adaptive LLM Agents from Persona-Driven Simulation with Preference-Batched GRPO. https://omanscience.com/en/articles/pb-grpo-learning-socially-adaptive-llm-agents-from-persona-driven-simulation-with-preference-batched-grpo
IEEE
J. Wang, J. Yin, X. Han, Y. Mei, J. Hao, and B. Guo, "PB-GRPO: Learning Socially Adaptive LLM Agents from Persona-Driven Simulation with Preference-Batched GRPO," https://omanscience.com/en/articles/pb-grpo-learning-socially-adaptive-llm-agents-from-persona-driven-simulation-with-preference-batched-grpo.