17 citations · 18 across the 6 of their papers we have counts for
5 papers · 1 filter
Multi-Source Transformer Architectures for Audiovisual Scene Classification
Wim Boes, Hugo Van hamme
In this technical report, the systems we submitted for subtask 1B of the DCASE 2021 challenge, regarding audiovisual scene classification, are described in detail. They are essenti…
Optimizing Temporal Resolution Of Convolutional Recurrent Neural Networks For Sound Event Detection
Wim Boes, Hugo Van hamme
In this technical report, the systems we submitted for subtask 4 of the DCASE 2021 challenge, regarding sound event detection, are described in detail. These models are closely rel…
Impact of temporal resolution on convolutional recurrent networks for audio tagging and sound event detection
Wim Boes, Hugo Van hamme
Many state-of-the-art systems for audio tagging and sound event detection employ convolutional recurrent neural architectures. Typically, they are trained in a mean teacher setting…
Multi-encoder attention-based architectures for sound recognition with partial visual assistance
Wim Boes, Hugo Van hamme
Large-scale sound recognition data sets typically consist of acoustic recordings obtained from multimedia libraries. As a consequence, modalities other than audio can often be expl…
Audiovisual Transformer Architectures for Large-Scale Classification and Synchronization of Weakly Labeled Audio Events
Wim Boes, Hugo Van hamme
We tackle the task of environmental event classification by drawing inspiration from the transformer neural network architecture used in machine translation. We modify this attenti…