activity
20242026
collaborators

6 papers

cs.CL2026

Evaluating Language Model Agency through Negotiations

Tim R. Davidson, Veniamin Veselovsky, Martin Josifoski +4

We introduce an approach to evaluate language model (LM) agency using negotiation games. This approach better reflects real-world use cases and addresses some of the shortcomings o…

cs.CV2025

Through the Theory of Mind's Eye: Reading Minds with Multimodal Video Large Language Models

Zhawnen Chen, Tianchun Wang, Yizhou Wang +4

Can large multimodal models have a human-like ability for emotional and social reasoning, and if so, how does it work? Recent research has discovered emergent theory-of-mind (ToM)…

cs.CL2024

Evaluating Large Language Models in Theory of Mind Tasks

Michal Kosinski

Eleven Large Language Models (LLMs) were assessed using a custom-made battery of false-belief tasks, considered a gold standard in testing Theory of Mind (ToM) in humans. The batte…

cs.MA2024

Crowd IQ -- Aggregating Opinions to Boost Performance

Michal Kosinski, Yoram Bachrach, Thore Graepel +2

We show how the quality of decisions based on the aggregated opinions of the crowd can be conveniently studied using a sample of individual responses to a standard IQ questionnaire…

cs.CV2024

Facial Width-to-Height Ratio Does Not Predict Self-Reported Behavioral Tendencies

Michal Kosinski

A growing number of studies have linked facial width-to-height ratio (fWHR) with various antisocial or violent behavioral tendencies. However, those studies have predominantly been…

cs.CV2024

Facial recognition technology and human raters can predict political orientation from images of expressionless faces even when controlling for demographics and self-presentation

Michal Kosinski, Poruz Khambatta, Yilun Wang

Carefully standardized facial images of 591 participants were taken in the laboratory, while controlling for self-presentation, facial expression, head orientation, and image prope…