859 citations · 1.5k across the 32 of their papers we have counts for
34 papers · 1 filter
Multi-label Zero-Shot Audio Classification with Temporal Attention
Duygu Dogan, Huang Xie, Toni Heittola +1
Zero-shot learning models are capable of classifying new classes by transferring knowledge from the seen classes using auxiliary information. While most of the existing zero-shot l…
Learning Contextual Tag Embeddings for Cross-Modal Alignment of Audio and Tags
Xavier Favory, Konstantinos Drossos, Tuomas Virtanen +1
Self-supervised audio representation learning offers an attractive alternative for obtaining generic audio embeddings, capable to be employed into various downstream tasks. Publish…
Robust Audio-Based Vehicle Counting in Low-to-Moderate Traffic Flow
Slobodan Djukanović, Jiři Matas, Tuomas Virtanen
The paper presents a method for audio-based vehicle counting (VC) in low-to-moderate traffic using one-channel sound. We formulate VC as a regression problem, i.e., we predict the…
WaveTransformer: A Novel Architecture for Audio Captioning Based on Learning Temporal and Time-Frequency Information
An Tran, Konstantinos Drossos, Tuomas Virtanen
Automated audio captioning (AAC) is a novel task, where a method takes as an input an audio sample and outputs a textual description (i.e. a caption) of its contents. Most AAC meth…
Neural Network-based Acoustic Vehicle Counting
Slobodan Djukanović, Yash Patel, Jiři Matas +1
This paper addresses acoustic vehicle counting using one-channel audio. We predict the pass-by instants of vehicles from local minima of clipped vehicle-to-microphone distance. Thi…
Conditioned Time-Dilated Convolutions for Sound Event Detection
Konstantinos Drossos, Stylianos I. Mimilakis, Tuomas Virtanen
Sound event detection (SED) is the task of identifying sound events along with their onset and offset times. A recent, convolutional neural networks based SED method, proposed the…