5 papers
Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding
Lucy Lin, Ayush Jain, Yifan Liu +1
Large Multimodal Models (LMMs) have achieved remarkable success on images and short videos, yet scaling them to long videos remains challenging due to frame-centric tokenization an…
Pi-HOC: Pairwise 3D Human-Object Contact Estimation
Sravan Chittupalli, Ayush Jain, Dong Huang
Resolving real-world human-object interactions in images is a many-to-many challenge, in which disentangling fine-grained concurrent physical contact is particularly difficult. Exi…
From Thousands to Billions: 3D Visual Language Grounding via Render-Supervised Distillation from 2D VLMs
Ang Cao, Sergio Arnaud, Oleksandr Maksymets +12
3D vision-language grounding faces a fundamental data bottleneck: while 2D models train on billions of images, 3D models have access to only thousands of labeled scenes--a six-orde…
Unifying 2D and 3D Vision-Language Understanding
Ayush Jain, Alexander Swerdlow, Yuzhou Wang +5
Progress in 3D vision-language learning has been hindered by the scarcity of large-scale 3D datasets. We introduce UniVLG, a unified architecture for 2D and 3D vision-language unde…
Locate 3D: Real-World Object Localization via Self-Supervised Learning in 3D
Sergio Arnaud, Paul McVay, Ada Martin +19
We present LOCATE 3D, a model for localizing objects in 3D scenes from referring expressions like "the small coffee table between the sofa and the lamp." LOCATE 3D sets a new state…