activity
20172022
most citedA Transformer-based joint-encoding for Emotion Recognition and Sentiment Analysis

105 citations · 143 across the 11 of their papers we have counts for

collaborators

14 papers

cs.CV20222 cited

Transformers and CNNs both Beat Humans on SBIR

Omar Seddati, Stéphane Dupont, Saïd Mahmoudi +1

Sketch-based image retrieval (SBIR) is the task of retrieving natural images (photos) that match the semantics and the spatial configuration of hand-drawn sketch queries. The unive…

cs.CV20221 cited

Analysis of Co-Laughter Gesture Relationship on RGB videos in Dyadic Conversation Contex

Hugo Bohy, Ahmad Hammoudeh, Antoine Maiorca +2

The development of virtual agents has enabled human-avatar interactions to become increasingly rich and varied. Moreover, an expressive virtual agent i.e. that mimics the natural e…

cs.CV2021

Multi-level Attention Fusion Network for Audio-visual Event Recognition

Mathilde Brousmiche, Jean Rouat, Stéphane Dupont

Event classification is inherently sequential and multimodal. Therefore, deep neural models need to dynamically focus on the most relevant time window and/or modality of a video. I…

cs.CV20202 cited

Improved Soccer Action Spotting using both Audio and Video Streams

Bastien Vanderplaetse, Stéphane Dupont

In this paper, we propose a study on multi-modal (audio and video) action spotting and classification in soccer videos. Action spotting and classification are the tasks that consis…

cs.CL2020

Modulated Fusion using Transformer for Linguistic-Acoustic Emotion Recognition

Jean-Benoit Delbrouck, Noé Tits, Stéphane Dupont

This paper aims to bring a new lightweight yet powerful solution for the task of Emotion Recognition and Sentiment Analysis. Our motivation is to propose two architectures based on…

cs.IR2020

AVECL-UMONS database for audio-visual event classification and localization

Mathilde Brousmiche, Stéphane Dupont, Jean Rouat

We introduce the AVECL-UMons dataset for audio-visual event classification and localization in the context of office environments. The audio-visual dataset is composed of 11 event…