activity
20182026
most citedOLR 2021 Challenge: Datasets, Rules and Baselines

7 citations · 14 across the 31 of their papers we have counts for

collaborators
Showing eess.ASShow all

13 papers · 1 filter

eess.AS2026

HoliDubber: Holistic Video Dubbing for Complex Acoustic Scenes via Text-Guided Audio Synthesis

Wenhao Guan, Yifan Duan, Junxi Liu +6

Video dubbing is a cornerstone of multimedia content creation, aiming to synthesize synchronized acoustic sequences for visual streams. While Text-to-Speech (TTS) and Text-to-Audio…

eess.AS2025

SyncVoice: Simple and Effective Automatic Video Dubbing with Vision-Augmented TTS

Kaidi Wang, Yi He, Wenhao Guan +9

Automatic video dubbing aims to generate high-fidelity speech that is temporally aligned with visual content. However, existing methods still suffer from limited speech naturalness…

eess.AS2025

Phoenix-VAD: Streaming Semantic Endpoint Detection for Full-Duplex Speech Interaction

Weijie Wu, Wenhao Guan, Kaidi Wang +6

Spoken dialogue models have significantly advanced intelligent human-computer interaction, yet they lack a plug-and-play full-duplex prediction module for semantic endpoint detecti…

eess.AS2024

LAFMA: A Latent Flow Matching Model for Text-to-Audio Generation

Wenhao Guan, Kaidi Wang, Wangjin Zhou +6

Recently, the application of diffusion models has facilitated the significant development of speech and audio generation. Nevertheless, the quality of samples generated by diffusio…

eess.AS20241 cited

MM-TTS: Multi-modal Prompt based Style Transfer for Expressive Text-to-Speech Synthesis

Wenhao Guan, Yishuang Li, Tao Li +6

The style transfer task in Text-to-Speech refers to the process of transferring style information into text content to generate corresponding speech with a specific style. However,…

eess.AS2023

Community Detection Graph Convolutional Network for Overlap-Aware Speaker Diarization

Jie Wang, Zhicong Chen, Haodong Zhou +2

The clustering algorithm plays a crucial role in speaker diarization systems. However, traditional clustering algorithms suffer from the complex distribution of speaker embeddings…