activity
20162023
most citedGenerating Synthetic Data for Text Recognition

35 citations · 53 across the 9 of their papers we have counts for

collaborators

9 papers

cs.CV2023

United We Stand, Divided We Fall: UnityGraph for Unsupervised Procedure Learning from Videos

Siddhant Bansal, Chetan Arora, C. V. Jawahar

Given multiple videos of the same task, procedure learning addresses identifying the key-steps and determining their order to perform the task. For this purpose, existing approache…

cs.CV2023

Explaining Deep Face Algorithms through Visualization: A Survey

Thrupthi Ann John, Vineeth N Balasubramanian, C. V. Jawahar

Although current deep models for face tasks surpass human performance on some benchmarks, we do not understand how they work. Thus, we cannot predict how it will react to novel inp…

cs.CV2023

Understanding Video Scenes through Text: Insights from Text-based Video Question Answering

Soumya Jahagirdar, Minesh Mathew, Dimosthenis Karatzas +1

Researchers have extensively studied the field of vision and language, discovering that both visual and textual content is crucial for understanding scenes effectively. Particularl…

cs.CV2023

Towards Real-Time Analysis of Broadcast Badminton Videos

Nitin Nilesh, Tushar Sharma, Anurag Ghosh +1

Analysis of player movements is a crucial subset of sports analysis. Existing player movement analysis methods use recorded videos after the match is over. In this work, we propose…

cs.CV20231 cited

CueCAn: Cue Driven Contextual Attention For Identifying Missing Traffic Signs on Unconstrained Roads

Varun Gupta, Anbumani Subramanian, C. V. Jawahar +1

Unconstrained Asian roads often involve poor infrastructure, affecting overall road safety. Missing traffic signs are a regular part of such roads. Missing or non-existing object d…

cs.CV202214 cited

Lip-to-Speech Synthesis for Arbitrary Speakers in the Wild

Sindhu B Hegde, K R Prajwal, Rudrabha Mukhopadhyay +2

In this work, we address the problem of generating speech from silent lip videos for any speaker in the wild. In stark contrast to previous works, our method (i) is not restricted…