3 papers
cs.RO2026
Binary Tracking for Spatial QA and Navigation with Open Vision-Language Models
Dongbin Na, Chanwoo Kim, Soonbin Rho +3
This work addresses spatial question answering for service robots traversing long egocentric routes. Given a query such as "where can I find a dry cleaner on the way back home?", t…
cs.CV2026
Semantic Flip: Synthetic OOD Generation for Robust Refusal in Embodied Question Answering and Spatial Localization
Dongbin Na, Chanwoo Kim, Giyun Choi +1
Detecting unanswerable user queries remains essential for the reliable deployment of real-world embodied agents. However, modern vision-language models (VLMs) often generate overly…
cs.CV2025
Leveraging 2D Masked Reconstruction for Domain Adaptation of 3D Pose Estimation
Hansoo Park, Chanwoo Kim, Jihyeon Kim +4
RGB-based 3D pose estimation methods have been successful with the development of deep learning and the emergence of high-quality 3D pose datasets. However, most existing methods d…