2 citations · 3 across the 5 of their papers we have counts for
11 papers
Towards zero-shot Text-based voice editing using acoustic context conditioning, utterance embeddings, and reference encoders
Jason Fong, Yun Wang, Prabhav Agrawal +4
Text-based voice editing (TBVE) uses synthetic output from text-to-speech (TTS) systems to replace words in an original recording. Recent work has used neural models to produce edi…
GCT: Gated Contextual Transformer for Sequential Audio Tagging
Yuanbo Hou, Yun Wang, Wenwu Wang +1
Audio tagging aims to assign predefined tags to audio clips to indicate the class information of audio events. Sequential audio tagging (SAT) means detecting both the class informa…
CT-SAT: Contextual Transformer for Sequential Audio Tagging
Yuanbo Hou, Zhaoyi Liu, Bo Kang +2
Sequential audio event tagging can provide not only the type information of audio events, but also the order information between events and the number of events that occur in an au…
Transferring Voice Knowledge for Acoustic Event Detection: An Empirical Study
Dawei Liang, Yangyang Shi, Yun Wang +6
Detection of common events and scenes from audio is useful for extracting and understanding human contexts in daily life. Prior studies have shown that leveraging knowledge from a…
Do sound event representations generalize to other audio tasks? A case study in audio transfer learning
Anurag Kumar, Yun Wang, Vamsi Krishna Ithapu +1
Transfer learning is critical for efficient information transfer across multiple related learning problems. A simple, yet effective transfer learning approach utilizes deep neural…
Dual Application of Speech Enhancement for Automatic Speech Recognition
Ashutosh Pandey, Chunxi Liu, Yun Wang +1
In this work, we exploit speech enhancement for improving a recurrent neural network transducer (RNN-T) based ASR system. We employ a dense convolutional recurrent network (DCRN) f…