2 papers
cs.CV2026
Where to Look Matters: Learning Influential Views for VLM-based 3D Visual Grounding
Tsung-Chih Chiang, Hsuan-Kung Yang, Jou-Min Liu +4
Recent zero-shot 3D visual grounding methods leverage vision-language models (VLMs) to localize objects in 3D scenes from natural language queries. However, these methods typically…
cs.CV2024
Reprojection Errors as Prompts for Efficient Scene Coordinate Regression
Ting-Ru Liu, Hsuan-Kung Yang, Jou-Min Liu +5
Scene coordinate regression (SCR) methods have emerged as a promising area of research due to their potential for accurate visual localization. However, many existing SCR approache…