3 papers
cs.CV2026
Where to Look Matters: Learning Influential Views for VLM-based 3D Visual Grounding
Tsung-Chih Chiang, Hsuan-Kung Yang, Jou-Min Liu +4
Recent zero-shot 3D visual grounding methods leverage vision-language models (VLMs) to localize objects in 3D scenes from natural language queries. However, these methods typically…
cs.CV2024
Reprojection Errors as Prompts for Efficient Scene Coordinate Regression
Ting-Ru Liu, Hsuan-Kung Yang, Jou-Min Liu +5
Scene coordinate regression (SCR) methods have emerged as a promising area of research due to their potential for accurate visual localization. However, many existing SCR approache…
cs.RO2023
Visual Forecasting as a Mid-level Representation for Avoidance
Hsuan-Kung Yang, Tsung-Chih Chiang, Ting-Ru Liu +3
The challenge of navigation in environments with dynamic objects continues to be a central issue in the study of autonomous agents. While predictive methods hold promise, their rel…