papers

Publications (41)

eess.AS2024

MM-TTS: Multi-modal Prompt based Style Transfer for Expressive Text-to-Speech Synthesis

Wenhao Guan, Yishuang Li, Tao Li +6

The style transfer task in Text-to-Speech refers to the process of transferring style information into text content to generate corresponding speech with a specific style. However,…

eess.AS2023

Community Detection Graph Convolutional Network for Overlap-Aware Speaker Diarization

Jie Wang, Zhicong Chen, Haodong Zhou +2

The clustering algorithm plays a crucial role in speaker diarization systems. However, traditional clustering algorithms suffer from the complex distribution of speaker embeddings…

cs.CL2024

Enhancing Code-Switching Speech Recognition with LID-Based Collaborative Mixture of Experts Model

Hukai Huang, Jiayan Lin, Kaidi Wang +4

Due to the inherent difficulty in modeling phonetic similarities across different languages, code-switching speech recognition presents a formidable challenge. This study proposes…

eess.AS2026

HoliDubber: Holistic Video Dubbing for Complex Acoustic Scenes via Text-Guided Audio Synthesis

Wenhao Guan, Yifan Duan, Junxi Liu +6

Video dubbing is a cornerstone of multimedia content creation, aiming to synthesize synchronized acoustic sequences for visual streams. While Text-to-Speech (TTS) and Text-to-Audio…

cs.SD2025

Cross-attention and Self-attention for Audio-visual Speaker Diarization in MISP-Meeting Challenge

Zhaoyang Li, Haodong Zhou, Longjie Luo +4

This paper presents the system developed for Task 1 of the Multi-modal Information-based Speech Processing (MISP) 2025 Challenge. We introduce CASA-Net, an embedding fusion method…

cs.SD2025

DS-Codec: Dual-Stage Training with Mirror-to-NonMirror Architecture Switching for Speech Codec

Peijie Chen, Wenhao Guan, Kaidi Wang +4

Neural speech codecs are essential for advancing text-to-speech (TTS) systems. With the recent success of large language models in text generation, developing high-quality speech t…