activity
20192025
most citedMultimodal Fusion with Deep Neural Networks for Audio-Video Emotion Recognition

43 citations · 72 across the 14 of their papers we have counts for

collaborators
Showing cs.LGShow all

8 papers · 1 filter

cs.LG20222 cited

RCC-GAN: Regularized Compound Conditional GAN for Large-Scale Tabular Data Synthesis

Mohammad Esmaeilpour, Nourhene Chaalia, Adel Abusitta +3

This paper introduces a novel generative adversarial network (GAN) for synthesizing large-scale tabular databases which contain various features such as continuous, discrete, and b…

cs.LG2020

Improving Stability of LS-GANs for Audio and Speech Signals

Mohammad Esmaeilpour, Raymel Alfonso Sallo, Olivier St-Georges +2

In this paper we address the instability issue of generative adversarial network (GAN) by proposing a new similarity metric in unitary space of Schur decomposition for 2D represent…

cs.LG2019

Detection of Adversarial Attacks and Characterization of Adversarial Subspace

Mohammad Esmaeilpour, Patrick Cardinal, Alessandro Lameiras Koerich

Adversarial attacks have always been a serious threat for any data-driven model. In this paper, we explore subspaces of adversarial examples in unitary vector domain, and we propos…

cs.LG201918 cited

Emotion Recognition with Spatial Attention and Temporal Softmax Pooling

Masih Aminbeidokhti, Marco Pedersoli, Patrick Cardinal +1

Video-based emotion recognition is a challenging task because it requires to distinguish the small deformations of the human face that represent emotions, while being invariant to…

cs.LG2019

Universal Adversarial Audio Perturbations

Sajjad Abdoli, Luiz G. Hafemann, Jerome Rony +3

We demonstrate the existence of universal adversarial perturbations, which can fool a family of audio classification architectures, for both targeted and untargeted attack scenario…

cs.LG2019

Emotion Recognition Using Fusion of Audio and Video Features

Juan D. S. Ortega, Patrick Cardinal, Alessandro L. Koerich

In this paper we propose a fusion approach to continuous emotion recognition that combines visual and auditory modalities in their representation spaces to predict the arousal and…