5 citations · 13 across the 7 of their papers we have counts for
4 papers · 1 filter
Understanding Depth and Height Perception in Large Visual-Language Models
Shehreen Azad, Yash Jain, Rishit Garg +2
Geometric understanding - including depth and height perception - is fundamental to intelligence and crucial for navigating our environment. Despite the impressive capabilities of…
PEEKABOO: Interactive Video Generation via Masked-Diffusion
Yash Jain, Anshul Nasery, Vibhav Vineet +1
Modern video generation models like Sora have achieved remarkable success in producing high-quality videos. However, a significant limitation is their inability to offer interactiv…
DAMEX: Dataset-aware Mixture-of-Experts for visual understanding of mixture-of-datasets
Yash Jain, Harkirat Behl, Zsolt Kira +1
Construction of a universal detector poses a crucial question: How can we most effectively train a model on a large mixture of datasets? The answer lies in learning dataset-specifi…
Fine-grained Human Activity Recognition Using Virtual On-body Acceleration Data
Zikang Leng, Yash Jain, Hyeokhyen Kwon +1
Previous work has demonstrated that virtual accelerometry data, extracted from videos using cross-modality transfer approaches like IMUTube, is beneficial for training complex and…