5 papers
SelectTSL: Prompt-Guided Selective Target Sound Localization in Complex Scenarios
Ziyang Jiang, Yu Chen, Zexu Pan +5
Humans can selectively attend to a target sound and estimate its direction in complex scenarios, whereas such selective localization remains challenging for current deep learning-b…
pTSE-T: Presentation Target Speaker Extraction using Unaligned Text Cues
Ziyang Jiang, Jiahe Lei, Xueyan Chen +4
Target Speaker Extraction (TSE) aims to extract the clean speech of the target speaker in an audio mixture, eliminating irrelevant background noise and speech. While prior work has…
I2TTS: Image-indicated Immersive Text-to-speech Synthesis with Spatial Perception
Jiawei Zhang, Tian-Hao Zhang, Jun Wang +3
Controlling the style and characteristics of speech synthesis is crucial for adapting the output to specific contexts and user requirements. Previous Text-to-speech (TTS) works hav…
FaceSpeak: Expressive and High-Quality Speech Synthesis from Human Portraits of Different Styles
Tian-Hao Zhang, Jiawei Zhang, Jun Wang +2
Humans can perceive speakers' characteristics (e.g., identity, gender, personality and emotion) by their appearance, which are generally aligned to their voice style. Recently, vis…
Breaking Through the Spike: Spike Window Decoding for Accelerated and Precise Automatic Speech Recognition
Wei Zhang, Tian-Hao Zhang, Chao Luo +4
Recently, end-to-end automatic speech recognition has become the mainstream approach in both industry and academia. To optimize system performance in specific scenarios, the Weight…