19 papers
Bagpiper: Solving Open-Ended Audio Tasks via Rich Captions
Jinchuan Tian, Haoran Wang, Bo-Hao Su +14
Current audio foundation models typically rely on rigid, task-specific supervision (e.g., speech recognition), addressing isolated factors of audio rather than the whole. In contra…
ESPnet3: Infrastructure for Scalable Speech and Audio Research in the Foundation Model Era
Masao Someki, Alexander Polok, Carlos Carvalho +14
Recent speech research involves increasingly large datasets, complex models, and diverse experimental workflows. However, existing frameworks require substantial engineering effort…
Full-Duplex-Bench-v2: A Multi-Turn Evaluation Framework for Duplex Dialogue Systems with an Automated Examiner
Guan-Ting Lin, Shih-Yun Shan Kuan, Jiatong Shi +4
While full-duplex speech agents enable natural, low-latency interaction by speaking and listening simultaneously, their consistency and task performance in multi-turn settings rema…
Reasoning Beyond Majority Vote: An Explainable SpeechLM Framework for Speech Emotion Recognition
Bo-Hao Su, Hui-Ying Shih, Jinchuan Tian +4
Speech Emotion Recognition (SER) is typically trained and evaluated on majority-voted labels, which simplifies benchmarking but masks subjectivity and provides little transparency…
Optimizing Conversational Quality in Spoken Dialogue Systems with Reinforcement Learning from AI Feedback
Siddhant Arora, Jinchuan Tian, Jiatong Shi +4
Reinforcement learning from human or AI feedback (RLHF/RLAIF) for speech-in/speech-out dialogue systems (SDS) remains underexplored, with prior work largely limited to single seman…
Do Neural Codecs Generalize? A Controlled Study Across Unseen Languages and Non-Speech Tasks
Shih-Heng Wang, Jiatong Shi, Jinchuan Tian +2
This paper investigates three crucial yet underexplored aspects of the generalization capabilities of neural audio codecs (NACs): (i) whether NACs can generalize to unseen language…