activity
20242026
collaborators

30 papers

eess.AS2026

An Efficient vLLM-Based Inference Pipeline for Unified Audio Understanding and Generation

Haoran Wang, Jinchuan Tian, Siddhant Arora +1

While Large Multimodal Models excel in comprehension, high-throughput inference engines lack native support for multimodal generation. This is severe in Speech Language Models, whe…

cs.CL2026

Bagpiper: Solving Open-Ended Audio Tasks via Rich Captions

Jinchuan Tian, Haoran Wang, Bo-Hao Su +14

Current audio foundation models typically rely on rigid, task-specific supervision (e.g., speech recognition), addressing isolated factors of audio rather than the whole. In contra…

eess.AS2026

ESPnet3: Infrastructure for Scalable Speech and Audio Research in the Foundation Model Era

Masao Someki, Alexander Polok, Carlos Carvalho +14

Recent speech research involves increasingly large datasets, complex models, and diverse experimental workflows. However, existing frameworks require substantial engineering effort…

cs.SD2026

Bagpiper-Edit: Zero-Shot Open-Ended Audio Editing via Rich-Caption

Xun Gong, Jinchuan Tian, Haoran Wang +3

Current text-guided audio editing methods rely on paired training data, predefined operation templates, and separate processing pipelines across speech, music, and sound. We presen…

cs.SD2026

Speech-Hands: A Self-Reflection Voice Agentic Approach to Speech Recognition and Audio Reasoning with Omni Perception

Zhen Wan, Chao-Han Huck Yang, Jinchuan Tian +15

We introduce a voice-agentic framework that learns one critical omni-understanding skill: knowing when to trust itself versus when to consult external audio perception. Our work is…

eess.AS2026

Reasoning Beyond Majority Vote: An Explainable SpeechLM Framework for Speech Emotion Recognition

Bo-Hao Su, Hui-Ying Shih, Jinchuan Tian +4

Speech Emotion Recognition (SER) is typically trained and evaluated on majority-voted labels, which simplifies benchmarking but masks subjectivity and provides little transparency…