859 citations · 1.6k across the 52 of their papers we have counts for
17 papers · 1 filter
Zero-Shot Audio Classification with Factored Linear and Nonlinear Acoustic-Semantic Projections
Huang Xie, Okko Räsänen, Tuomas Virtanen
In this paper, we study zero-shot learning in audio classification through factored linear and nonlinear acoustic-semantic projections between audio instances and sound classes. Ze…
Zero-Shot Audio Classification via Semantic Embeddings
Huang Xie, Tuomas Virtanen
In this paper, we study zero-shot learning in audio classification via semantic embeddings extracted from textual labels and sentence descriptions of sound classes. Our goal is to…
A Curated Dataset of Urban Scenes for Audio-Visual Scene Analysis
Shanshan Wang, Annamaria Mesaros, Toni Heittola +1
This paper introduces a curated dataset of urban scenes for audio-visual scene analysis which consists of carefully selected and recorded material. The data was recorded in multipl…
Learning Contextual Tag Embeddings for Cross-Modal Alignment of Audio and Tags
Xavier Favory, Konstantinos Drossos, Tuomas Virtanen +1
Self-supervised audio representation learning offers an attractive alternative for obtaining generic audio embeddings, capable to be employed into various downstream tasks. Publish…
Robust Audio-Based Vehicle Counting in Low-to-Moderate Traffic Flow
Slobodan Djukanović, Jiři Matas, Tuomas Virtanen
The paper presents a method for audio-based vehicle counting (VC) in low-to-moderate traffic using one-channel sound. We formulate VC as a regression problem, i.e., we predict the…
WaveTransformer: A Novel Architecture for Audio Captioning Based on Learning Temporal and Time-Frequency Information
An Tran, Konstantinos Drossos, Tuomas Virtanen
Automated audio captioning (AAC) is a novel task, where a method takes as an input an audio sample and outputs a textual description (i.e. a caption) of its contents. Most AAC meth…