3 papers
eess.AS2024
Separate Anything You Describe
Xubo Liu, Qiuqiang Kong, Yan Zhao +7
Language-queried audio source separation (LASS) is a new paradigm for computational auditory scene analysis (CASA). LASS aims to separate a target sound from an audio mixture given…
cs.SD2024
Joint Multi-scale Cross-lingual Speaking Style Transfer with Bidirectional Attention Mechanism for Automatic Dubbing
Jingbei Li, Sipan Li, Ping Chen +7
Automatic dubbing, which generates a corresponding version of the input speech in another language, could be widely utilized in many real-world scenarios such as video and game loc…
cs.SD2024
AudioLDM 2: Learning Holistic Audio Generation with Self-supervised Pretraining
Haohe Liu, Yi Yuan, Xubo Liu +7
Although audio generation shares commonalities across different types of audio, such as speech, music, and sound effects, designing models for each type requires careful considerat…