activity
20242026
collaborators

5 papers

cs.CL2026

ConvApparel: A Benchmark Dataset and Validation Framework for User Simulators in Conversational Recommenders

Ofer Meshi, Krisztian Balog, Sally Goldman +5

The promise of LLM-based user simulators to improve conversational AI is hindered by a critical "realism gap," leading to systems that are optimized for simulated interactions, but…

cs.CL2026

LLMs and people both learn to form conventions -- just not with each other

Cameron R. Jones, Agnese Lombardi, Kyle Mahowald +1

Humans align to one another in conversation -- adopting shared conventions that ease communication. We test whether LLMs form the same kinds of conventions in a multimodal communic…

cs.CL2025

Dissecting the Ullman Variations with a SCALPEL: Why do LLMs fail at Trivial Alterations to the False Belief Task?

Zhiqiang Pi, Annapurna Vadaparty, Benjamin K. Bergen +1

Recent empirical results have sparked a debate about whether or not Large Language Models (LLMs) are capable of Theory of Mind (ToM). While some have found LLMs to be successful on…

cs.CL2025

Large Language Models Pass the Turing Test

Cameron R. Jones, Benjamin K. Bergen

We evaluated 4 systems (ELIZA, GPT-4o, LLaMa-3.1-405B, and GPT-4.5) in two randomised, controlled, and pre-registered Turing tests on independent populations. Participants had 5 mi…

cs.CL2024

Lies, Damned Lies, and Distributional Language Statistics: Persuasion and Deception with Large Language Models

Cameron R. Jones, Benjamin K. Bergen

Large Language Models (LLMs) can generate content that is as persuasive as human-written text and appear capable of selectively producing deceptive outputs. These capabilities rais…