13 citations · 13 across the 4 of their papers we have counts for
4 papers · 1 filter
CoordRefer: Coordinate-Aware 3D Visual Grounding from Multiview Images
Haijie Li, Jiaxin Zhang, Dave Zhenyu Chen +3
Multiview image-based 3D visual grounding predicts a coordinate frame to define a coordinate system and then regresses a 3D bounding box for localization. However, existing methods…
Breaking the Horizontal Prior: From Long-Tailed Orientation Bias to Roll-Robust Monocular Depth Estimation
Kaihua Tang, Ziqing Xia, Xiaoxu Zheng +4
Despite recent advances in Monocular Depth Estimation, state-of-the-art depth foundation models remain vulnerable to robustness issues. Particularly, even slight camera rolls can r…
Generating Context-Aware Natural Answers for Questions in 3D Scenes
Mohammed Munzer Dwedari, Matthias Niessner, Dave Zhenyu Chen
3D question answering is a young field in 3D vision-language that is yet to be explored. Previous methods are limited to a pre-defined answer space and cannot generate answers natu…
Text2Tex: Text-driven Texture Synthesis via Diffusion Models
Dave Zhenyu Chen, Yawar Siddiqui, Hsin-Ying Lee +2
We present Text2Tex, a novel method for generating high-quality textures for 3D meshes from the given text prompts. Our method incorporates inpainting into a pre-trained depth-awar…