activity
20172024
most citedDeep Learning for Audio Signal Processing

859 citations · 1.6k across the 52 of their papers we have counts for

collaborators
Showing 2020Show all

17 papers · 1 filter

eess.AS2020★ 2 cited

Zero-Shot Audio Classification with Factored Linear and Nonlinear Acoustic-Semantic Projections

Huang Xie, Okko Räsänen, Tuomas Virtanen

In this paper, we study zero-shot learning in audio classification through factored linear and nonlinear acoustic-semantic projections between audio instances and sound classes. Ze…

eess.AS2020★ 4 cited

Zero-Shot Audio Classification via Semantic Embeddings

Huang Xie, Tuomas Virtanen

In this paper, we study zero-shot learning in audio classification via semantic embeddings extracted from textual labels and sentence descriptions of sound classes. Our goal is to…

eess.AS2020★ 3 cited

A Curated Dataset of Urban Scenes for Audio-Visual Scene Analysis

Shanshan Wang, Annamaria Mesaros, Toni Heittola +1

This paper introduces a curated dataset of urban scenes for audio-visual scene analysis which consists of carefully selected and recorded material. The data was recorded in multipl…

cs.SD2020

Learning Contextual Tag Embeddings for Cross-Modal Alignment of Audio and Tags

Xavier Favory, Konstantinos Drossos, Tuomas Virtanen +1

Self-supervised audio representation learning offers an attractive alternative for obtaining generic audio embeddings, capable to be employed into various downstream tasks. Publish…

cs.SD2020

Robust Audio-Based Vehicle Counting in Low-to-Moderate Traffic Flow

Slobodan Djukanović, Jiři Matas, Tuomas Virtanen

The paper presents a method for audio-based vehicle counting (VC) in low-to-moderate traffic using one-channel sound. We formulate VC as a regression problem, i.e., we predict the…

cs.SD2020

WaveTransformer: A Novel Architecture for Audio Captioning Based on Learning Temporal and Time-Frequency Information

An Tran, Konstantinos Drossos, Tuomas Virtanen

Automated audio captioning (AAC) is a novel task, where a method takes as an input an audio sample and outputs a textual description (i.e. a caption) of its contents. Most AAC meth…