45 citations · 45 across the 3 of their papers we have counts for
5 papers
Crowdsourcing a Dataset of Audio Captions
Samuel Lipping, Konstantinos Drossos, Tuomas Virtanen
Audio captioning is a novel field of multi-modal translation and it is the task of creating a textual description of the content of an audio signal (e.g. "people talking in a big r…
Stacked Convolutional and Recurrent Neural Networks for Music Emotion Recognition
Miroslav Malik, Sharath Adavanne, Konstantinos Drossos +3
This paper studies the emotion recognition from musical tracks in the 2-dimensional valence-arousal (V-A) emotional space. We propose a method based on convolutional (CNN) and recu…
Stacked Convolutional and Recurrent Neural Networks for Bird Audio Detection
Sharath Adavanne, Konstantinos Drossos, Emre Çakır +1
This paper studies the detection of bird calls in audio segments using stacked convolutional and recurrent neural networks. Data augmentation by blocks mixing and domain adaptation…
Automated Audio Captioning with Recurrent Neural Networks
Konstantinos Drossos, Sharath Adavanne, Tuomas Virtanen
We present the first approach to automated audio captioning. We employ an encoder-decoder scheme with an alignment model in between. The input to the encoder is a sequence of log m…
Convolutional Recurrent Neural Networks for Bird Audio Detection
EmreÇakır, Sharath Adavanne, Giambattista Parascandolo +2
Bird sounds possess distinctive spectral structure which may exhibit small shifts in spectrum depending on the bird species and environmental conditions. In this paper, we propose…