10 papers
FaciliTrain: Practicing Facilitation Skills through AI-Simulated Group Dialogue
Hang Jiang, Yuanxin Zhu, Diyi Yang +3
Skilled facilitation supports inclusive small-group dialogue, but deliberate practice is hard to scale: it depends on expert coaches, live practice partners, and iterative feedback…
Rating the Pitch, Not the Product: User Evaluations of LLMs Reflect Expectations More Than Performance
Robert Morabito, Tyler McDonald, Charitra Viswanath +4
Imagine two users interact with the same LLM. One has been told it is the cutting-edge flagship model; the other, an older, weaker model. They walk away with markedly different rat…
DART: Mitigating Harm Drift in Difference-Aware LLMs via Distill-Audit-Repair Training
Ziwen Pan, Zihan Liang, Jad Kabbara +1
Large language models (LLMs) tuned for safety often avoid acknowledging demographic differences, even when such acknowledgment is factually correct (e.g., ancestry-based disease in…
Common to Whom? Regional Cultural Commonsense and LLM Bias in India
Sangmitra Madhusudan, Trush Shashank More, Steph Buongiorno +3
Existing cultural commonsense benchmarks treat nations as monolithic, assuming uniform practices within national boundaries. But does cultural commonsense hold uniformly within a n…
LLM Targeted Underperformance Disproportionately Impacts Vulnerable Users
Elinor Poole-Dayan, Deb Roy, Jad Kabbara
While state-of-the-art large language models (LLMs) have shown impressive performance on many tasks, there has been extensive research on undesirable model behavior such as halluci…
Computational Analysis of Conversation Dynamics through Participant Responsivity
Margaret Hughes, Brandon Roy, Elinor Poole-Dayan +2
Growing literature explores toxicity and polarization in discourse, with comparatively less work on characterizing what makes dialogue prosocial and constructive. We explore conver…