most citedUsing Active Speaker Faces for Diarization in TV shows

2 citations · 5 across the 6 of their papers we have counts for

collaborators

6 papers

cs.MM2022

Audio-Visual Activity Guided Cross-Modal Identity Association for Active Speaker Detection

Rahul Sharma, Shrikanth Narayanan

Active speaker detection in videos addresses associating a source face, visible in the video frames, with the underlying speech in the audio modality. The two primary sources of in…

eess.IV20221 cited

Unsupervised active speaker detection in media content using cross-modal information

Rahul Sharma, Shrikanth Narayanan

We present a cross-modal unsupervised framework for active speaker detection in media content such as TV shows and movies. Machine learning advances have enabled impressive perform…

cs.LG2022

Federated Learning with Noisy User Feedback

Rahul Sharma, Anil Ramakrishna, Ansel MacLaughlin +5

Machine Learning (ML) systems are getting increasingly popular, and drive more and more applications and services in our daily life. This has led to growing concerns over user priv…

cs.MM20222 cited

Using Active Speaker Faces for Diarization in TV shows

Rahul Sharma, Shrikanth Narayanan

Speaker diarization is one of the critical components of computational media intelligence as it enables a character-level analysis of story portrayals and media content understandi…

cs.CV20222 cited

Audio visual character profiles for detecting background characters in entertainment media

Rahul Sharma, Shrikanth Narayanan

An essential goal of computational media intelligence is to support understanding how media stories -- be it news, commercial or entertainment media -- represent and reflect societ…

cs.MM2017

Multichannel Attention Network for Analyzing Visual Behavior in Public Speaking

Rahul Sharma, Tanaya Guha, Gaurav Sharma

Public speaking is an important aspect of human communication and interaction. The majority of computational work on public speaking concentrates on analyzing the spoken content, a…