most citedSelf-Supervised learning with cross-modal transformers for emotion recognition

40 citations · 50 across the 3 of their papers we have counts for

collaborators

6 papers

cs.CV2021

Audiovisual Highlight Detection in Videos

Karel Mundnich, Alexandra Fenster, Aparna Khare +1

In this paper, we test the hypothesis that interesting events in unstructured videos are inherently audiovisual. We combine deep image representations for object recognition and sc…

cs.CL202040 cited

Self-Supervised learning with cross-modal transformers for emotion recognition

Aparna Khare, Srinivas Parthasarathy, Shiva Sundaram

Emotion recognition is a challenging task due to limited availability of in-the-wild labeled datasets. Self-supervised learning has shown improvements on tasks with limited labeled…

cs.CL2020

Multi-modal embeddings using multi-task learning for emotion recognition

Aparna Khare, Srinivas Parthasarathy, Shiva Sundaram

General embeddings like word2vec, GloVe and ELMo have shown a lot of success in natural language tasks. The embeddings are typically extracted from models that are built on general…

eess.AS20207 cited

Multiresolution and Multimodal Speech Recognition with Transformers

Georgios Paraskevopoulos, Srinivas Parthasarathy, Aparna Khare +1

This paper presents an audio visual automatic speech recognition (AV-ASR) system using a Transformer-based architecture. We particularly focus on the scene context provided by the…

cs.SD2020

Fully Learnable Front-End for Multi-Channel Acoustic Modeling using Semi-Supervised Learning

Sanna Wager, Aparna Khare, Minhua Wu +2

In this work, we investigated the teacher-student training paradigm to train a fully learnable multi-channel acoustic model for far-field automatic speech recognition (ASR). Using…

cs.SD20203 cited

Multi-channel Acoustic Modeling using Mixed Bitrate OPUS Compression

Aparna Khare, Shiva Sundaram, Minhua Wu

Recent literature has shown that a learned front end with multi-channel audio input can outperform traditional beam-forming algorithms for automatic speech recognition (ASR). In th…