1 citations · 1 across the 1 of their papers we have counts for
Showing cs.CVShow all
2 papers · 1 filter
cs.CV2025
HoneyBee: Data Recipes for Vision-Language Reasoners
Hritik Bansal, Devendra Singh Sachan, Kai-Wei Chang +4
Recent advances in vision-language models (VLMs) have made them highly effective at reasoning tasks. However, the principles underlying the construction of performant VL reasoning…
cs.CV2025
VideoPhy-2: A Challenging Action-Centric Physical Commonsense Evaluation in Video Generation
Hritik Bansal, Clark Peng, Yonatan Bitton +3
Large-scale video generative models, capable of creating realistic videos of diverse visual concepts, are strong candidates for general-purpose physical world simulators. However,…