2 papers
cs.CV2026
Time Blindness: Why Video-Language Models Can't See What Humans Can?
Ujjwal Upadhyay, Mukul Ranjan, Zhiqiang Shen +1
Recent advances in vision-language models (VLMs) have made impressive strides in understanding spatio-temporal relationships in videos. However, when spatial information is obscure…
cs.CV2025
3DCoMPaT: An improved Large-scale 3D Vision Dataset for Compositional Recognition
Habib Slim, Xiang Li, Yuchen Li +8
In this work, we present 3DCoMPaT, a multimodal 2D/3D dataset with 160 million rendered views of more than 10 million stylized 3D shapes carefully annotated at the part-inst…