40 citations · 50 across the 5 of their papers we have counts for
11 papers
Disentanglement for audio-visual emotion recognition using multitask setup
Raghuveer Peri, Srinivas Parthasarathy, Charles Bradshaw +1
Deep learning models trained on audio-visual data have been successfully used to achieve state-of-the-art performance for emotion recognition. In particular, models trained with mu…
Audiovisual Highlight Detection in Videos
Karel Mundnich, Alexandra Fenster, Aparna Khare +1
In this paper, we test the hypothesis that interesting events in unstructured videos are inherently audiovisual. We combine deep image representations for object recognition and sc…
Detecting expressions with multimodal transformers
Srinivas Parthasarathy, Shiva Sundaram
Developing machine learning algorithms to understand person-to-person engagement can result in natural user experiences for communal devices such as Amazon Alexa. Among other cues…
Training Strategies to Handle Missing Modalities for Audio-Visual Expression Recognition
Srinivas Parthasarathy, Shiva Sundaram
Automatic audio-visual expression recognition can play an important role in communication services such as tele-health, VOIP calls and human-machine interaction. Accuracy of audio-…
Self-Supervised learning with cross-modal transformers for emotion recognition
Aparna Khare, Srinivas Parthasarathy, Shiva Sundaram
Emotion recognition is a challenging task due to limited availability of in-the-wild labeled datasets. Self-supervised learning has shown improvements on tasks with limited labeled…
Multi-modal embeddings using multi-task learning for emotion recognition
Aparna Khare, Srinivas Parthasarathy, Shiva Sundaram
General embeddings like word2vec, GloVe and ELMo have shown a lot of success in natural language tasks. The embeddings are typically extracted from models that are built on general…