Abstract

Social preferences can promote cooperation in multi-agent reinforcement learning, but existing approaches often require agents to observe the rewards of their peers. In many real-world interactions, however, an agent can, as humans do, observe others' behavior and outcomes without access to their private reward signals. We introduce self-referenced social preferences, in which each agent learns a model of its own reward, applies it to other agents' observed transitions to assess their outcomes from its own perspective, and feeds these self-referenced assessments into standard social preferences. We study two ways to incorporate these assessments: modifying the learning reward, or using them to weight policy updates. We evaluate the approach on three sequential social dilemmas, Escape Room, Clean Up, and Commons Harvest, which require volunteering, public-good contribution, and resource restraint, respectively. Across all three environments, agents learn cooperative behavior without observing others' rewards, including in settings where independent learners fail to cooperate, and frequently achieve more equitable divisions of jointly produced returns than agents with access to true rewards. The effective integration point depends on the social preference: inequity aversion works best in the reward together with a value look-ahead, whereas a purely benevolent preference benefits from policy-update weighting. Under partial observability, the policy-update approach continues to support cooperation. These results show that explicit access to other agents' reward signals is not necessary for learning cooperative behavior: social preferences can instead be grounded in self-referenced assessments of others' outcomes derived from their observed behavior.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Mohamed, M. A., Kotamreddy, H., & Jose, M. M. (2026). Self-Referenced Social Preferences: Cooperation without Observing Others Rewards. https://omanscience.com/en/articles/self-referenced-social-preferences-cooperation-without-observing-others-rewards

MLA 9

Mohamed, Mohamed Ayman, et al. "Self-Referenced Social Preferences: Cooperation without Observing Others Rewards." https://omanscience.com/en/articles/self-referenced-social-preferences-cooperation-without-observing-others-rewards.

Chicago (author–date)

Mohamed, Mohamed Ayman, Harshil Kotamreddy, and Marcos Menon Jose. 2026. "Self-Referenced Social Preferences: Cooperation without Observing Others Rewards." https://omanscience.com/en/articles/self-referenced-social-preferences-cooperation-without-observing-others-rewards.

Harvard

Mohamed, M. A., Kotamreddy, H. and Jose, M. M. (2026) 'Self-Referenced Social Preferences: Cooperation without Observing Others Rewards', Available at: https://omanscience.com/en/articles/self-referenced-social-preferences-cooperation-without-observing-others-rewards.

Vancouver

Mohamed MA, Kotamreddy H, Jose MM. Self-Referenced Social Preferences: Cooperation without Observing Others Rewards. https://omanscience.com/en/articles/self-referenced-social-preferences-cooperation-without-observing-others-rewards

IEEE

M. A. Mohamed, H. Kotamreddy, and M. M. Jose, "Self-Referenced Social Preferences: Cooperation without Observing Others Rewards," https://omanscience.com/en/articles/self-referenced-social-preferences-cooperation-without-observing-others-rewards.