2 citations · 5 across the 6 of their papers we have counts for
6 papers
Audio-Visual Activity Guided Cross-Modal Identity Association for Active Speaker Detection
Rahul Sharma, Shrikanth Narayanan
Active speaker detection in videos addresses associating a source face, visible in the video frames, with the underlying speech in the audio modality. The two primary sources of in…
Unsupervised active speaker detection in media content using cross-modal information
Rahul Sharma, Shrikanth Narayanan
We present a cross-modal unsupervised framework for active speaker detection in media content such as TV shows and movies. Machine learning advances have enabled impressive perform…
Federated Learning with Noisy User Feedback
Rahul Sharma, Anil Ramakrishna, Ansel MacLaughlin +5
Machine Learning (ML) systems are getting increasingly popular, and drive more and more applications and services in our daily life. This has led to growing concerns over user priv…
Using Active Speaker Faces for Diarization in TV shows
Rahul Sharma, Shrikanth Narayanan
Speaker diarization is one of the critical components of computational media intelligence as it enables a character-level analysis of story portrayals and media content understandi…
Audio visual character profiles for detecting background characters in entertainment media
Rahul Sharma, Shrikanth Narayanan
An essential goal of computational media intelligence is to support understanding how media stories -- be it news, commercial or entertainment media -- represent and reflect societ…
Multichannel Attention Network for Analyzing Visual Behavior in Public Speaking
Rahul Sharma, Tanaya Guha, Gaurav Sharma
Public speaking is an important aspect of human communication and interaction. The majority of computational work on public speaking concentrates on analyzing the spoken content, a…