11 citations · 13 across the 6 of their papers we have counts for
6 papers
LanSER: Language-Model Supported Speech Emotion Recognition
Taesik Gong, Josh Belanich, Krishna Somandepalli +3
Speech emotion recognition (SER) models typically rely on costly human-labeled data for training, making scaling methods to large speech datasets and nuanced emotion taxonomies dif…
MM-AU:Towards Multimodal Understanding of Advertisement Videos
Digbalay Bose, Rajat Hebbar, Tiantian Feng +3
Advertisement videos (ads) play an integral part in the domain of Internet e-commerce as they amplify the reach of particular products to a broad audience or can serve as a medium…
Contextually-rich human affect perception using multimodal scene information
Digbalay Bose, Rajat Hebbar, Krishna Somandepalli +1
The process of human affect understanding involves the ability to infer person specific emotional states from various sources including images, speech, and language. Affect percept…
Heterogeneous Graph Learning for Acoustic Event Classification
Amir Shirian, Mona Ahmadian, Krishna Somandepalli +1
Heterogeneous graphs provide a compact, efficient, and scalable way to model data involving multiple disparate modalities. This makes modeling audiovisual data using heterogeneous…
A dataset for Audio-Visual Sound Event Detection in Movies
Rajat Hebbar, Digbalay Bose, Krishna Somandepalli +2
Audio event detection is a widely studied audio processing task, with applications ranging from self-driving cars to healthcare. In-the-wild datasets such as Audioset have propelle…
Visually-aware Acoustic Event Detection using Heterogeneous Graphs
Amir Shirian, Krishna Somandepalli, Victor Sanchez +1
Perception of auditory events is inherently multimodal relying on both audio and visual cues. A large number of existing multimodal approaches process each modality using modality-…