collaborators

6 papers

eess.AS2026

An Efficient vLLM-Based Inference Pipeline for Unified Audio Understanding and Generation

Haoran Wang, Jinchuan Tian, Siddhant Arora +1

While Large Multimodal Models excel in comprehension, high-throughput inference engines lack native support for multimodal generation. This is severe in Speech Language Models, whe…

cs.CL2026

Bagpiper-TTS: Natural Language Guided Universal Speech Synthesis

Jinchuan Tian, Haoran Wang, Siddhant Arora +6

Classical TTS systems typically rely on rigid input formats and predefined metadata slots, limiting their ability to fulfill flexible user requirements. This paper introduces Bagpi…

cs.CV2026

Interest Entanglement: The Hidden Barrier to Blind Super-Resolution Optimization

Junxiong Lin, Xinji Mai, Qianyu Guo +5

Fidelity and perceptual quality are two inherently competing and conflicting objectives in the image super-resolution (SR) task. Different loss functions focus on these objectives…

cs.CV2026

Customizing Video Portraits via Identity-ActionDecoupling

Junxiong Lin, Haoran Wang, Xinji Mai +4

Identity-Preserving Text-to-Video Generation (IPT2V) seeks to synthesize a temporally coherent video from a reference image and a textual description, while simultaneously preservi…

cs.SD2026

Bagpiper-Edit: Zero-Shot Open-Ended Audio Editing via Rich-Caption

Xun Gong, Jinchuan Tian, Haoran Wang +3

Current text-guided audio editing methods rely on paired training data, predefined operation templates, and separate processing pipelines across speech, music, and sound. We presen…

cs.SD2025

PURE Codec: Progressive Unfolding of Residual Entropy for Speech Codec Learning

Jiatong Shi, Haoran Wang, William Chen +4

Neural speech codecs have achieved strong performance in low-bitrate compression, but residual vector quantization (RVQ) often suffers from unstable training and ineffective decomp…