36 citations · 179 across the 24 of their papers we have counts for
26 papers · 1 filter
GCT: Gated Contextual Transformer for Sequential Audio Tagging
Yuanbo Hou, Yun Wang, Wenwu Wang +1
Audio tagging aims to assign predefined tags to audio clips to indicate the class information of audio events. Sequential audio tagging (SAT) means detecting both the class informa…
Automated Audio Captioning via Fusion of Low- and High- Dimensional Features
Jianyuan Sun, Xubo Liu, Xinhao Mei +3
Automated audio captioning (AAC) aims to describe the content of an audio clip using simple sentences. Existing AAC methods are developed based on an encoder-decoder architecture t…
RaDur: A Reference-aware and Duration-robust Network for Target Sound Detection
Dongchao Yang, Helin Wang, Zhongjie Ye +2
Target sound detection (TSD) aims to detect the target sound from a mixture audio given the reference information. Previous methods use a conditional network to extract a sound-dis…
Time-domain Speech Enhancement with Generative Adversarial Learning
Feiyang Xiao, Jian Guan, Qiuqiang Kong +1
Speech enhancement aims to obtain speech signals with high intelligibility and quality from noisy speech. Recent work has demonstrated the excellent performance of time-domain deep…
Enhancing Audio Augmentation Methods with Consistency Learning
Turab Iqbal, Karim Helwani, Arvindh Krishnaswamy +1
Data augmentation is an inexpensive way to increase training data diversity and is commonly achieved via transformations of existing data. For tasks such as classification, there i…
An Improved Event-Independent Network for Polyphonic Sound Event Localization and Detection
Yin Cao, Turab Iqbal, Qiuqiang Kong +3
Polyphonic sound event localization and detection (SELD), which jointly performs sound event detection (SED) and direction-of-arrival (DoA) estimation, detects the type and occurre…