6 papers
DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation
Wei Zhou, Wanyi Ning, Yinshang Guo +3
Audio-visual speech enhancement under real-world conditions remains challenging due to unreliable visual inputs and the lack of large-scale training data with realistic acoustic co…
PS4: Proxy-Supervised Joint Training for Real Target Speaker Extraction
Wanyi Ning, Wei Zhou, Yingpeng Li +3
Training target speaker extraction (TSE) models for real conversational mixtures remains challenging because large-scale training corpora and clean target speech for supervision ar…
Latent Speech-Text Transformer
Yen-Ju Lu, Yashesh Gaur, Wei Zhou +8
Auto-regressive speech-text models pre-trained on interleaved text tokens and discretized speech tokens demonstrate strong speech understanding and generation, yet remain substanti…
Can Speech LLMs Think while Listening?
Yi-Jen Shih, Desh Raj, Chunyang Wu +6
Recent advances in speech large language models (speech LLMs) have enabled seamless spoken interactions, but these systems still struggle with complex reasoning tasks. Previously,…
Frozen Large Language Models Can Perceive Paralinguistic Aspects of Speech
Wonjune Kang, Junteng Jia, Chunyang Wu +8
This work studies the capabilities of a large language model (LLM) to understand paralinguistic aspects of speech without fine-tuning its weights. We utilize an end-to-end system w…
CJST: CTC Compressor based Joint Speech and Text Training for Decoder-Only ASR
Wei Zhou, Junteng Jia, Leda Sari +2
CTC compressor can be an effective approach to integrate audio encoders to decoder-only models, which has gained growing interest for different speech applications. In this work, w…