SocialReasoning-Bench: Measuring Whether AI Agents Act in Users’ Best Interests
Tyler Payne, Will Epperson, Safoora Yousefi, Zachary Huang, Gagan Bansal, Wenyue Hua, Maya Murad, Asli Celikyilmaz, Saleema Amershi
Summary
AI agents are moving into social contexts. When agents manage calendars, negotiate purchases, or interact with other agents on a user’s behalf, they need more than task competence: they need social reasoning, an understanding of what the user wants, what the counterparty wants, and what information to reveal, protect, or push back on. SocialReasoning-Bench evaluates that ability. The benchmark tests whether an agent can negotiate for a user in two realistic settings, Calendar Coordination and Marketplace Negotiation, against a counterparty with its own goals, private information, and sometimes adversarial intent. It measures both outcomes and process, scoring agents on outcome optimality (how much value they secure for the user) and due diligence (whether they follow a competent decision-making process).
Current frontier models often leave value on the table. They usually complete the task, but they frequently accept suboptimal meeting times or poor deals instead of advocating effectively for the user. Prompting helps, but it is not enough: even with explicit guidance to act in the user’s best interest, performance remains well below what a trustworthy delegate should achieve.
Citation
SocialReasoning-Bench: Measuring Whether AI Agents Act in Users’ Best Interests
Tyler Payne, Will Epperson, Safoora Yousefi, Zachary Huang, Gagan Bansal, Wenyue Hua, Maya Murad, Asli Celikyilmaz, Saleema Amershi
A benchmark for measuring AI agents’ negotiation ability in calendar coordination and marketplace settings.
Microsoft Research Blog. 2026.
Project Blog Code
Will Epperson