5 papers
Hear2Act: Benchmarking When Prosody Should Change What an Assistant Does
Xinyi Liu, Hooshang Nayyeri, Dilek Hakkani-Tur +6
Prosodic cues can convey task-relevant information that alters the trajectory and outcome of a task-oriented dialogue, even when the words themselves remain unchanged. Yet existing…
Beyond Individual Personas: Aligning Synthetic Dialogue to Population-Level Behavior Distributions
Xinyi Liu, Rinat Khaziev, Hooshang Nayyeri +3
Synthetic dialogue corpora are increasingly used as proxies for target dialogue data, yet persona-grounded generators optimize individual conversations rather than corpus compositi…
Towards Self-Improving Error Diagnosis in Multi-Agent Systems
Jiazheng Li, Emine Yilmaz, Bei Chen +1
Large Language Model (LLM)-based Multi-Agent Systems (MAS) enable complex problem-solving but introduce significant debugging challenges, characterized by long interaction traces,…
ATOD: An Evaluation Framework and Benchmark for Agentic Task-Oriented Dialogue Systems
Yifei Zhang, Hooshang Nayyeri, Rinat Khaziev +4
Recent advances in task-oriented dialogue (TOD) systems, driven by large language models (LLMs) with extensive API and tool integration, have enabled conversational agents to coord…
Adaptive Multi-Agent Response Refinement in Conversational Systems
Soyeong Jeong, Aparna Elangovan, Emine Yilmaz +1
Large Language Models (LLMs) have demonstrated remarkable success in conversational systems by generating human-like responses. However, they can fall short, especially when requir…