collaborators

15 papers

eess.AS2026

An Efficient vLLM-Based Inference Pipeline for Unified Audio Understanding and Generation

Haoran Wang, Jinchuan Tian, Siddhant Arora +1

While Large Multimodal Models excel in comprehension, high-throughput inference engines lack native support for multimodal generation. This is severe in Speech Language Models, whe…

eess.AS2026

Comparative Reasoning: Making an Audio Language Model Better at Comparing Emotions

Abinay Reddy Naini, Jaeyeon Kim, Chao-Han Huck Yang +2

Large audio-language models (LALMs) can reason about audio, yet it remains unclear whether they can perform comparative judgments between two speech signals along emotional, enviro…

cs.CL2026

Bagpiper-TTS: Natural Language Guided Universal Speech Synthesis

Jinchuan Tian, Haoran Wang, Siddhant Arora +6

Classical TTS systems typically rely on rigid input formats and predefined metadata slots, limiting their ability to fulfill flexible user requirements. This paper introduces Bagpi…

eess.AS2026

ESPnet3: Infrastructure for Scalable Speech and Audio Research in the Foundation Model Era

Masao Someki, Alexander Polok, Carlos Carvalho +14

Recent speech research involves increasingly large datasets, complex models, and diverse experimental workflows. However, existing frameworks require substantial engineering effort…

cs.SD2026

Online Predictive Coding for Dual-Mode Self-Supervised Speech Model

Keita Goto, Takashi Maekaku, Jin Sakuma +3

Dual-mode self-supervised speech models are pre-trained to handle streaming and non-streaming conditions simultaneously. However, their attention is computed over different context…

cs.SD2026

Bagpiper-Edit: Zero-Shot Open-Ended Audio Editing via Rich-Caption

Xun Gong, Jinchuan Tian, Haoran Wang +3

Current text-guided audio editing methods rely on paired training data, predefined operation templates, and separate processing pipelines across speech, music, and sound. We presen…