activity
20182021
most citedMultimodal Speech Emotion Recognition and Ambiguity Resolution

42 citations · 47 across the 3 of their papers we have counts for

collaborators

7 papers

cs.SD20211 cited

LyricJam: A system for generating lyrics for live instrumental music

Olga Vechtomova, Gaurav Sahu, Dhruv Kumar

We describe a real-time system that receives a live audio stream from a jam session and generates lyric lines that are congruent with the live music being played. Two novel approac…

cs.AI2021

Towards A Multi-agent System for Online Hate Speech Detection

Gaurav Sahu, Robin Cohen, Olga Vechtomova

This paper envisions a multi-agent system for detecting the presence of hate speech in online social media platforms such as Twitter and Facebook. We introduce a novel framework em…

cs.CL20204 cited

Generation of lyrics lines conditioned on music audio clips

Olga Vechtomova, Gaurav Sahu, Dhruv Kumar

We present a system for generating novel lyrics lines conditioned on music audio. A bimodal neural network model learns to generate lines conditioned on any given short audio clip.…

cs.CL2019

Adaptive Fusion Techniques for Multimodal Data

Gaurav Sahu, Olga Vechtomova

Effective fusion of data from multiple modalities, such as video, speech, and text, is challenging due to the heterogeneous nature of multimodal data. In this paper, we propose ada…

cs.CL2019

Adversarial Learning on the Latent Space for Diverse Dialog Generation

Kashif Khan, Gaurav Sahu, Vikash Balasubramanian +2

Generating relevant responses in a dialog is challenging, and requires not only proper modeling of context in the conversation but also being able to generate fluent sentences duri…

cs.LG201942 cited

Multimodal Speech Emotion Recognition and Ambiguity Resolution

Gaurav Sahu

Identifying emotion from speech is a non-trivial task pertaining to the ambiguous definition of emotion itself. In this work, we adopt a feature-engineering based approach to tackl…