18 citations · 18 across the 4 of their papers we have counts for
4 papers · 1 filter
Multi-task Regularization Based on Infrequent Classes for Audio Captioning
Emre Çakır, Konstantinos Drossos, Tuomas Virtanen
Audio captioning is a multi-modal task, focusing on using natural language for describing the contents of general audio. Most audio captioning methods are based on deep neural netw…
End-to-End Polyphonic Sound Event Detection Using Convolutional Recurrent Neural Networks with Learned Time-Frequency Representation Input
Emre Çakır, Tuomas Virtanen
Sound event detection systems typically consist of two stages: extracting hand-crafted features from the raw audio waveform, and learning a mapping between these features and the t…
Stacked Convolutional and Recurrent Neural Networks for Bird Audio Detection
Sharath Adavanne, Konstantinos Drossos, Emre Çakır +1
This paper studies the detection of bird calls in audio segments using stacked convolutional and recurrent neural networks. Data augmentation by blocks mixing and domain adaptation…
Convolutional Recurrent Neural Networks for Bird Audio Detection
EmreÇakır, Sharath Adavanne, Giambattista Parascandolo +2
Bird sounds possess distinctive spectral structure which may exhibit small shifts in spectrum depending on the bird species and environmental conditions. In this paper, we propose…