35 citations · 53 across the 9 of their papers we have counts for
9 papers
United We Stand, Divided We Fall: UnityGraph for Unsupervised Procedure Learning from Videos
Siddhant Bansal, Chetan Arora, C. V. Jawahar
Given multiple videos of the same task, procedure learning addresses identifying the key-steps and determining their order to perform the task. For this purpose, existing approache…
Explaining Deep Face Algorithms through Visualization: A Survey
Thrupthi Ann John, Vineeth N Balasubramanian, C. V. Jawahar
Although current deep models for face tasks surpass human performance on some benchmarks, we do not understand how they work. Thus, we cannot predict how it will react to novel inp…
Understanding Video Scenes through Text: Insights from Text-based Video Question Answering
Soumya Jahagirdar, Minesh Mathew, Dimosthenis Karatzas +1
Researchers have extensively studied the field of vision and language, discovering that both visual and textual content is crucial for understanding scenes effectively. Particularl…
Towards Real-Time Analysis of Broadcast Badminton Videos
Nitin Nilesh, Tushar Sharma, Anurag Ghosh +1
Analysis of player movements is a crucial subset of sports analysis. Existing player movement analysis methods use recorded videos after the match is over. In this work, we propose…
CueCAn: Cue Driven Contextual Attention For Identifying Missing Traffic Signs on Unconstrained Roads
Varun Gupta, Anbumani Subramanian, C. V. Jawahar +1
Unconstrained Asian roads often involve poor infrastructure, affecting overall road safety. Missing traffic signs are a regular part of such roads. Missing or non-existing object d…
Lip-to-Speech Synthesis for Arbitrary Speakers in the Wild
Sindhu B Hegde, K R Prajwal, Rudrabha Mukhopadhyay +2
In this work, we address the problem of generating speech from silent lip videos for any speaker in the wild. In stark contrast to previous works, our method (i) is not restricted…