2 citations · 7 across the 33 of their papers we have counts for
36 papers
An Efficient vLLM-Based Inference Pipeline for Unified Audio Understanding and Generation
Haoran Wang, Jinchuan Tian, Siddhant Arora +1
While Large Multimodal Models excel in comprehension, high-throughput inference engines lack native support for multimodal generation. This is severe in Speech Language Models, whe…
ESPnet3: Infrastructure for Scalable Speech and Audio Research in the Foundation Model Era
Masao Someki, Alexander Polok, Carlos Carvalho +14
Recent speech research involves increasingly large datasets, complex models, and diverse experimental workflows. However, existing frameworks require substantial engineering effort…
Bagpiper-Edit: Zero-Shot Open-Ended Audio Editing via Rich-Caption
Xun Gong, Jinchuan Tian, Haoran Wang +3
Current text-guided audio editing methods rely on paired training data, predefined operation templates, and separate processing pipelines across speech, music, and sound. We presen…
Bagpiper: Solving Open-Ended Audio Tasks via Rich Captions
Jinchuan Tian, Haoran Wang, Bo-Hao Su +14
Current audio foundation models typically rely on rigid, task-specific supervision (e.g., speech recognition), addressing isolated factors of audio rather than the whole. In contra…
Optimizing Conversational Quality in Spoken Dialogue Systems with Reinforcement Learning from AI Feedback
Siddhant Arora, Jinchuan Tian, Jiatong Shi +4
Reinforcement learning from human or AI feedback (RLHF/RLAIF) for speech-in/speech-out dialogue systems (SDS) remains underexplored, with prior work largely limited to single seman…
Do Neural Codecs Generalize? A Controlled Study Across Unseen Languages and Non-Speech Tasks
Shih-Heng Wang, Jiatong Shi, Jinchuan Tian +2
This paper investigates three crucial yet underexplored aspects of the generalization capabilities of neural audio codecs (NACs): (i) whether NACs can generalize to unseen language…