Showing cs.CVShow all
2 papers · 1 filter
cs.CV2025
3DCoMPaT: An improved Large-scale 3D Vision Dataset for Compositional Recognition
Habib Slim, Xiang Li, Yuchen Li +8
In this work, we present 3DCoMPaT, a multimodal 2D/3D dataset with 160 million rendered views of more than 10 million stylized 3D shapes carefully annotated at the part-inst…
cs.CV2024
CoT3DRef: Chain-of-Thoughts Data-Efficient 3D Visual Grounding
Eslam Abdelrahman, Mohamed Ayman, Mahmoud Ahmed +2
3D visual grounding is the ability to localize objects in 3D scenes conditioned by utterances. Most existing methods devote the referring head to localize the referred object direc…