34 citations · 177 across the 55 of their papers we have counts for
4 papers · 2 filters
Language-specific Acoustic Boundary Learning for Mandarin-English Code-switching Speech Recognition
Zhiyun Fan, Linhao Dong, Chen Shen +4
Code-switching speech recognition (CSSR) transcribes speech that switches between multiple languages or dialects within a single sentence. The main challenge in this task is that d…
Make-An-Audio 2: Temporal-Enhanced Text-to-Audio Generation
Jiawei Huang, Yi Ren, Rongjie Huang +7
Large diffusion models have been successful in text-to-audio (T2A) synthesis tasks, but they often suffer from common issues such as semantic misalignment and poor temporal consist…
StyleS2ST: Zero-shot Style Transfer for Direct Speech-to-speech Translation
Kun Song, Yi Ren, Yi Lei +5
Direct speech-to-speech translation (S2ST) has gradually become popular as it has many advantages compared with cascade S2ST. However, current research mainly focuses on the accura…
ByteCover3: Accurate Cover Song Identification on Short Queries
Xingjian Du, Zijie Wang, Xia Liang +3
Deep learning based methods have become a paradigm for cover song identification (CSI) in recent years, where the ByteCover systems have achieved state-of-the-art results on all th…