3 citations · 3 across the 4 of their papers we have counts for
4 papers
HOI-Ref: Hand-Object Interaction Referral in Egocentric Vision
Siddhant Bansal, Michael Wray, Dima Damen
Large Vision Language Models (VLMs) are now the de facto state-of-the-art for a number of tasks including visual question answering, recognising objects, and spatial referral. In t…
United We Stand, Divided We Fall: UnityGraph for Unsupervised Procedure Learning from Videos
Siddhant Bansal, Chetan Arora, C. V. Jawahar
Given multiple videos of the same task, procedure learning addresses identifying the key-steps and determining their order to perform the task. For this purpose, existing approache…
My View is the Best View: Procedure Learning from Egocentric Videos
Siddhant Bansal, Chetan Arora, C. V. Jawahar
Procedure learning involves identifying the key-steps and determining their logical order to perform a task. Existing approaches commonly use third-person videos for learning the p…
Making AI 'Smart': Bridging AI and Cognitive Science
Madhav Agarwal, Siddhant Bansal
The last two decades have seen tremendous advances in Artificial Intelligence. The exponential growth in terms of computation capabilities has given us hope of developing humans li…