17 citations · 18 across the 6 of their papers we have counts for
6 papers
Multi-Source Transformer Architectures for Audiovisual Scene Classification
Wim Boes, Hugo Van hamme
In this technical report, the systems we submitted for subtask 1B of the DCASE 2021 challenge, regarding audiovisual scene classification, are described in detail. They are essenti…
Optimizing Temporal Resolution Of Convolutional Recurrent Neural Networks For Sound Event Detection
Wim Boes, Hugo Van hamme
In this technical report, the systems we submitted for subtask 4 of the DCASE 2021 challenge, regarding sound event detection, are described in detail. These models are closely rel…
Impact of temporal resolution on convolutional recurrent networks for audio tagging and sound event detection
Wim Boes, Hugo Van hamme
Many state-of-the-art systems for audio tagging and sound event detection employ convolutional recurrent neural architectures. Typically, they are trained in a mean teacher setting…
Multi-encoder attention-based architectures for sound recognition with partial visual assistance
Wim Boes, Hugo Van hamme
Large-scale sound recognition data sets typically consist of acoustic recordings obtained from multimedia libraries. As a consequence, modalities other than audio can often be expl…
On the long-term learning ability of LSTM LMs
Wim Boes, Robbe Van Rompaey, Lyan Verwimp +3
We inspect the long-term learning ability of Long Short-Term Memory language models (LSTM LMs) by evaluating a contextual extension based on the Continuous Bag-of-Words (CBOW) mode…
Audiovisual Transformer Architectures for Large-Scale Classification and Synchronization of Weakly Labeled Audio Events
Wim Boes, Hugo Van hamme
We tackle the task of environmental event classification by drawing inspiration from the transformer neural network architecture used in machine translation. We modify this attenti…