Abstract

Building LLMs that behave well socially, not merely correctly, requires Building LLMs that behave well socially, not merely correctly, requires more than producing locally helpful responses. A socially competent agent must infer users' unstated goals, respect their preferences, and adapt as the conversation unfolds. These behaviors are inherently multi-turn and social, making them hard to optimize: real interaction data is scarce, and user preferences are typically latent rather than directly observable. To address these challenges, we build on a persona-driven social simulation environment (consisting of a persona library, LLM-based user simulators, and a user-satisfaction scoring system ranging from [0, 1]), to introduce preference-batched GRPO (PB-GRPO), a post-training algorithm that learns socially adaptive policies from conversation-level feedback. Compared to vanilla GRPO, PB-GRPO computes advantages using a normalization estimated across a bucket of users with similar preferences, stabilizing training across a diverse social population. Empirical evidence shows that PB-GRPO improves models' social behavior over strong reinforcement learning baselines in our simulated environment.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Wang, J., Yin, J., Han, X., Mei, Y., Hao, J., & Guo, B. (2026). PB-GRPO: Learning Socially Adaptive LLM Agents from Persona-Driven Simulation with Preference-Batched GRPO. https://omanscience.com/en/articles/pb-grpo-learning-socially-adaptive-llm-agents-from-persona-driven-simulation-with-preference-batched-grpo

MLA 9

Wang, Jingquan, et al. "PB-GRPO: Learning Socially Adaptive LLM Agents from Persona-Driven Simulation with Preference-Batched GRPO." https://omanscience.com/en/articles/pb-grpo-learning-socially-adaptive-llm-agents-from-persona-driven-simulation-with-preference-batched-grpo.

Chicago (author–date)

Wang, Jingquan, Jun Yin, Xu Han, Yongsheng Mei, Jie Hao, and Bin Guo. 2026. "PB-GRPO: Learning Socially Adaptive LLM Agents from Persona-Driven Simulation with Preference-Batched GRPO." https://omanscience.com/en/articles/pb-grpo-learning-socially-adaptive-llm-agents-from-persona-driven-simulation-with-preference-batched-grpo.

Harvard

Wang, J., Yin, J., Han, X., Mei, Y., Hao, J. and Guo, B. (2026) 'PB-GRPO: Learning Socially Adaptive LLM Agents from Persona-Driven Simulation with Preference-Batched GRPO', Available at: https://omanscience.com/en/articles/pb-grpo-learning-socially-adaptive-llm-agents-from-persona-driven-simulation-with-preference-batched-grpo.

Vancouver

Wang J, Yin J, Han X, Mei Y, Hao J, Guo B. PB-GRPO: Learning Socially Adaptive LLM Agents from Persona-Driven Simulation with Preference-Batched GRPO. https://omanscience.com/en/articles/pb-grpo-learning-socially-adaptive-llm-agents-from-persona-driven-simulation-with-preference-batched-grpo

IEEE

J. Wang, J. Yin, X. Han, Y. Mei, J. Hao, and B. Guo, "PB-GRPO: Learning Socially Adaptive LLM Agents from Persona-Driven Simulation with Preference-Batched GRPO," https://omanscience.com/en/articles/pb-grpo-learning-socially-adaptive-llm-agents-from-persona-driven-simulation-with-preference-batched-grpo.