activity
20182022
most citedCT-SAT: Contextual Transformer for Sequential Audio Tagging

2 citations · 3 across the 5 of their papers we have counts for

collaborators

11 papers

cs.SD2022

Towards zero-shot Text-based voice editing using acoustic context conditioning, utterance embeddings, and reference encoders

Jason Fong, Yun Wang, Prabhav Agrawal +4

Text-based voice editing (TBVE) uses synthetic output from text-to-speech (TTS) systems to replace words in an original recording. Recent work has used neural models to produce edi…

cs.SD2022

GCT: Gated Contextual Transformer for Sequential Audio Tagging

Yuanbo Hou, Yun Wang, Wenwu Wang +1

Audio tagging aims to assign predefined tags to audio clips to indicate the class information of audio events. Sequential audio tagging (SAT) means detecting both the class informa…

cs.SD20222 cited

CT-SAT: Contextual Transformer for Sequential Audio Tagging

Yuanbo Hou, Zhaoyi Liu, Bo Kang +2

Sequential audio event tagging can provide not only the type information of audio events, but also the order information between events and the number of events that occur in an au…

cs.SD20211 cited

Transferring Voice Knowledge for Acoustic Event Detection: An Empirical Study

Dawei Liang, Yangyang Shi, Yun Wang +6

Detection of common events and scenes from audio is useful for extracting and understanding human contexts in daily life. Prior studies have shown that leveraging knowledge from a…

cs.SD2021

Do sound event representations generalize to other audio tasks? A case study in audio transfer learning

Anurag Kumar, Yun Wang, Vamsi Krishna Ithapu +1

Transfer learning is critical for efficient information transfer across multiple related learning problems. A simple, yet effective transfer learning approach utilizes deep neural…

cs.SD2020

Dual Application of Speech Enhancement for Automatic Speech Recognition

Ashutosh Pandey, Chunxi Liu, Yun Wang +1

In this work, we exploit speech enhancement for improving a recurrent neural network transducer (RNN-T) based ASR system. We employ a dense convolutional recurrent network (DCRN) f…