5 papers
When Policies Change Probabilities: Modular Decision-Making for LLM Code Review
Rasvik Kudum, Max Corbett, Hitansh Paliwal +3
LLM code reviewers often estimate patch risk and make approval decisions in one prompt. A probability should depend on evidence; costs should determine the action taken from it. We…
EASE Configuration Facilitates A Reproducible Science of LLM Social Simulations
Sneheel Sarangi, Maximilian Puelma Touzel, Aurélien Bück-Kaeffer +3
LLMs are increasingly deployed to simulate social interactions, yet many of the existing simulators remain ad hoc and monolithic. This lack of architectural standardization prevent…
The Cookbook: Design Space of LLM-based Social Simulations
Aurélien Bück-Kaeffer, Sneheel Sarangi, Maximilian Puelma Touzel +3
Studies attempting to simulate human behavior with grow in numbers while LLM-only social networks have started appearing outside of controlled settings…
Evaluating Generalization Capabilities of LLM-Based Agents in Mixed-Motive Scenarios Using Concordia
Chandler Smith, Marwa Abdulhai, Manfred Diaz +83
Large Language Model (LLM) agents have demonstrated impressive capabilities for social interaction and are increasingly being deployed in situations where they might engage with bo…
Decompose-ToM: Enhancing Theory of Mind Reasoning in Large Language Models through Simulation and Task Decomposition
Sneheel Sarangi, Maha Elgarf, Hanan Salam
Theory of Mind (ToM) is the ability to understand and reflect on the mental states of others. Although this capability is crucial for human interaction, testing on Large Language M…