79 citations · 102 across the 4 of their papers we have counts for
8 papers
A Pre-trained Audio-Visual Transformer for Emotion Recognition
Minh Tran, Mohammad Soleymani
In this paper, we introduce a pretrained audio-visual Transformer trained on more than 500k utterances from nearly 4000 celebrities from the VoxCeleb2 dataset for human behavior un…
Speaker Turn Modeling for Dialogue Act Classification
Zihao He, Leili Tavabi, Kristina Lerman +1
Dialogue Act (DA) classification is the task of classifying utterances with respect to the function they serve in a dialogue. Existing approaches to DA classification model utteran…
Modeling Dynamics of Facial Behavior for Mental Health Assessment
Minh Tran, Ellen Bradley, Michelle Matvey +2
Facial action unit (FAU) intensities are popular descriptors for the analysis of facial behavior. However, FAUs are sparsely represented when only a few are activated at a time. In…
Affective Computing for Large-Scale Heterogeneous Multimedia Data: A Survey
Sicheng Zhao, Shangfei Wang, Mohammad Soleymani +2
The wide popularity of digital photography and social networks has generated a rapidly growing volume of multimedia data (i.e., image, music, and video), resulting in a great deman…
Polysemous Visual-Semantic Embedding for Cross-Modal Retrieval
Yale Song, Mohammad Soleymani
Visual-semantic embedding aims to find a shared latent space where related visual and textual instances are close to each other. Most current methods learn injective embedding func…
AVEC 2019 Workshop and Challenge: State-of-Mind, Detecting Depression with AI, and Cross-Cultural Affect Recognition
Fabien Ringeval, Björn Schuller, Michel Valstar +14
The Audio/Visual Emotion Challenge and Workshop (AVEC 2019) "State-of-Mind, Detecting Depression with AI, and Cross-cultural Affect Recognition" is the ninth competition event aime…