17 citations · 32 across the 8 of their papers we have counts for
3 papers · 1 filter
Deep neural network Based Low-latency Speech Separation with Asymmetric analysis-Synthesis Window Pair
Shanshan Wang, Gaurav Naithani, Archontis Politis +1
Time-frequency masking or spectrum prediction computed via short symmetric windows are commonly used in low-latency deep neural network (DNN) based source separation. In this paper…
Audio-visual scene classification: analysis of DCASE 2021 Challenge submissions
Shanshan Wang, Toni Heittola, Annamaria Mesaros +1
This paper presents the details of the Audio-Visual Scene Classification task in the DCASE 2021 Challenge (Task 1 Subtask B). The task is concerned with classification using audio…
A Curated Dataset of Urban Scenes for Audio-Visual Scene Analysis
Shanshan Wang, Annamaria Mesaros, Toni Heittola +1
This paper introduces a curated dataset of urban scenes for audio-visual scene analysis which consists of carefully selected and recorded material. The data was recorded in multipl…