265 citations · 532 across the 34 of their papers we have counts for
7 papers · 1 filter
An Encoder-Decoder Based Audio Captioning System With Transfer and Reinforcement Learning
Xinhao Mei, Qiushi Huang, Xubo Liu +10
Automated audio captioning aims to use natural language to describe the content of audio data. This paper presents an audio captioning system with an encoder-decoder architecture,…
Audio Captioning Transformer
Xinhao Mei, Xubo Liu, Qiushi Huang +2
Audio captioning aims to automatically generate a natural language description of an audio clip. Most captioning models follow an encoder-decoder architecture, where the decoder pr…
Sound Event Detection: A Tutorial
Annamaria Mesaros, Toni Heittola, Tuomas Virtanen +1
The goal of automatic sound event detection (SED) methods is to recognize what is happening in an audio signal and when it is happening. In practice, the goal is to recognize at wh…
Conditional Sound Generation Using Neural Discrete Time-Frequency Representation Learning
Xubo Liu, Turab Iqbal, Jinzheng Zhao +3
Deep generative models have recently achieved impressive performance in speech and music synthesis. However, compared to the generation of those domain-specific sounds, generating…
Event-Independent Network for Polyphonic Sound Event Localization and Detection
Yin Cao, Turab Iqbal, Qiuqiang Kong +3
Polyphonic sound event localization and detection is not only detecting what sound events are happening but localizing corresponding sound sources. This series of tasks was first i…
Sound Event Localization and Detection Using CRNN on Pairs of Microphones
Francois Grondin, James Glass, Iwona Sobieraj +1
This paper proposes sound event localization and detection methods from multichannel recording. The proposed system is based on two Convolutional Recurrent Neural Networks (CRNNs)…