1 citations · 1 across the 4 of their papers we have counts for
4 papers
BLR-MoE: Boosted Language-Routing Mixture of Experts for Domain-Robust Multilingual E2E ASR
Guodong Ma, Wenxuan Wang, Lifeng Zhou +3
Recently, the Mixture of Expert (MoE) architecture, such as LR-MoE, is often used to alleviate the impact of language confusion on the multilingual ASR (MASR) task. However, it sti…
Cross-Modal Denoising: A Novel Training Paradigm for Enhancing Speech-Image Retrieval
Lifeng Zhou, Yuke Li, Rui Deng +2
The success of speech-image retrieval relies on establishing an effective alignment between speech and image. Existing methods often model cross-modal interaction through simple co…
Coarse-to-fine Alignment Makes Better Speech-image Retrieval
Lifeng Zhou, Yuke Li
In this paper, we propose a novel framework for speech-image retrieval. We utilize speech-image contrastive (SIC) learning tasks to align speech and image representations at a coar…
Acoustic Pornography Recognition Using Convolutional Neural Networks and Bag of Refinements
Lifeng Zhou, Kaifeng Wei, Yuke Li +3
A large number of pornographic audios publicly available on the Internet seriously threaten the mental and physical health of children, but these audios are rarely detected and fil…