4 papers
TurnBench: A Multi-Domain Benchmark for Turn-Taking Dynamics in Spoken Dialogue
Freeman Jiang, Ramon Sanabria, Soham Deshmukh +16
Speakers in natural conversation take turns speaking and listening, deciding in real time when to take, hold, or yield the floor. However, turn-taking evaluation remains limited du…
CS-YODAS: A Mined Dataset of In-the-Wild Code-Switched Speech
Brian Yan, Qingzheng Wang, Matthew Wiesner +9
We present CS-YODAS, a Creative Commons-licensed dataset of in-the-wild code-switched speech mined from multilingual YouTube data. Code-switching (CS), or the alternation between l…
AudioChat: Unified Audio Storytelling, Editing, and Understanding with Transfusion Forcing
William Chen, Prem Seetharaman, Rithesh Kumar +4
Despite recent breakthroughs, audio foundation models struggle in processing complex multi-source acoustic scenes. We refer to this challenging domain as audio stories, which can h…
ADEPT: RL-Aligned Agentic Decoding of Emotion via Evidence Probing Tools -- From Consensus Learning to Ambiguity-Driven Emotion Reasoning
Esther Sun, Bo-Hao Su, Abinay Reddy Naini +2
Speech Large Language Models (SLLMs) enable high-level emotion reasoning but often produce ungrounded, text-biased judgments without verifiable acoustic evidence. In contrast, self…