papers

Publications (7)

cs.MA2026

Talk is Cheap, Communication is Hard: Dynamic Grounding Failures and Repair in Multi-Agent Negotiation

Yiheng Yao, Chelsea Zou, Robert D. Hawkins

Grounding is the collaborative process of establishing mutual belief sufficient for a communicative goal. While static grounding maps language to a shared context, dynamic groundin…

cs.CL2026

A Unified Definition of Hallucination: It's The World Model, Stupid!

Emmy Liu, Varun Gangal, Chelsea Zou +7

Despite numerous attempts at mitigation since the inception of language models, hallucinations remain a persistent problem even in today's frontier LLMs. Why is this? We review exi…

cs.CV2026

Abstracted Gaussian Prototypes for True One-Shot Concept Learning

Chelsea Zou, Kenneth J. Kurtz

We introduce a cluster-based generative image segmentation framework to encode higher-level representations of visual concepts based on one-shot learning inspired by the Omniglot C…

cs.RO2023

ARDIE: AR, Dialogue, and Eye Gaze Policies for Human-Robot Collaboration

Chelsea Zou, Kishan Chandan, Yan Ding +1

Human-robot collaboration (HRC) has become increasingly relevant in industrial, household, and commercial settings. However, the effectiveness of such collaborations is highly depe…

cs.LG2026

Humanity's Last Exam

Long Phan, Alice Gatti, Ziwen Han +1144

Benchmarks are important tools for tracking the rapid advancements in large language model (LLM) capabilities. However, benchmarks are not keeping pace in difficulty: LLMs now achi…

cs.MA2026

CalBench: Evaluating Coordination-Privacy Trade-offs in Multi-Agent LLMs

Chelsea Zou, Yiheng Yao, Selena She +2

Personal AI assistants are beginning to act as delegates with access to calendars, inboxes, and user preferences. Calendar scheduling makes the trust problem concrete: an assistant…

cs.AI2025

Thinking, Faithful and Stable: Mitigating Hallucinations in LLMs

Chelsea Zou, Yiheng Yao, Basant Khalil

This project develops a self correcting framework for large language models (LLMs) that detects and mitigates hallucinations during multi-step reasoning. Rather than relying solely…