Abstract
Developing socially intelligent AI remains heavily dependent on human-annotated data, limiting the scale and breadth of social understanding models can acquire. Methods that derive training signals from unlabeled data offer a path beyond this dependence, but social predictions lack the verification oracles available in mathematics and coding. Moreover, core social targets such as affect, intent, preference, and pragmatic meaning are often ambiguous. The same behavior can support multiple plausible interpretations, making it difficult to verify which is best supported. To address this challenge, we introduce Reinforcement Learning with Comparative Evidence (RLCE), a reinforcement learning method that learns social understanding from unlabeled training data without constructing rewards from ground-truth annotations. Given distinct answers in a rollout group, RLCE constructs evidence tests that identify observable evidence favoring an answer over another, validates these tests against the input sample, and aggregates test outcomes to determine the best-supported interpretation. Tests are regenerated as the policy produces new answers, enabling them to evolve with the policy. Across four benchmarks spanning affect, pragmatics, communicative intent, and preference, RLCE attains the strongest performance among seven methods that use no ground-truth training labels for rewards, including consensus, policy LLM-judge verification, multimodal co-evolution, and rubric-based rewards. Gains over the strongest baseline reach up to +18.93 points. Analyses further show that RLCE exhibits a larger share of reward variation between correct and incorrect predictions than compared rubric methods, can overturn erroneous policy-derived preferences, and benefits from pairwise test construction, compositional test aggregation, and on-policy test evolution.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Ong, K., Ryan, Y., Boughorbel, S., Necula, V., Shi, J. W. L., Mao, R., Lee, R. K. W., Kuek, A., Chen, N. F., Cambria, E., Mengaldo, G., & Liang, P. P. (2026). Reinforcement Learning with Comparative Evidence for Social Intelligence. https://omanscience.com/en/articles/reinforcement-learning-with-comparative-evidence-for-social-intelligence
MLA 9
Ong, Keane, et al. "Reinforcement Learning with Comparative Evidence for Social Intelligence." https://omanscience.com/en/articles/reinforcement-learning-with-comparative-evidence-for-social-intelligence.
Chicago (author–date)
Ong, Keane, Yuriel Ryan, Sabri Boughorbel, Vladimir Necula, Jack Wei Lun Shi, Rui Mao, Roy Ka-Wei Lee, Adriel Kuek, Nancy F. Chen, Erik Cambria, Gianmarco Mengaldo, and Paul Pu Liang. 2026. "Reinforcement Learning with Comparative Evidence for Social Intelligence." https://omanscience.com/en/articles/reinforcement-learning-with-comparative-evidence-for-social-intelligence.
Harvard
Ong, K., Ryan, Y., Boughorbel, S., Necula, V., Shi, J. W. L., Mao, R., Lee, R. K. W., Kuek, A., Chen, N. F., Cambria, E., Mengaldo, G. and Liang, P. P. (2026) 'Reinforcement Learning with Comparative Evidence for Social Intelligence', Available at: https://omanscience.com/en/articles/reinforcement-learning-with-comparative-evidence-for-social-intelligence.
Vancouver
Ong K, Ryan Y, Boughorbel S, Necula V, Shi JWL, Mao R, et al. Reinforcement Learning with Comparative Evidence for Social Intelligence. https://omanscience.com/en/articles/reinforcement-learning-with-comparative-evidence-for-social-intelligence
IEEE
K. Ong, Y. Ryan, S. Boughorbel, V. Necula, J. W. L. Shi, R. Mao, R. K. W. Lee, A. Kuek, N. F. Chen, E. Cambria, G. Mengaldo, and P. P. Liang, "Reinforcement Learning with Comparative Evidence for Social Intelligence," https://omanscience.com/en/articles/reinforcement-learning-with-comparative-evidence-for-social-intelligence.