8 papers
PanoGrounder: Bridging 2D and 3D with Panoramic Scene Representations for VLM-based 3D Visual Grounding
Seongmin Jung, Seongho Choi, Gunwoo Jeon +2
3D Visual Grounding (3DVG) is a critical bridge from vision-language perception to robotics, requiring both language understanding and 3D scene reasoning. Traditional supervised mo…
Vision-aligned Latent Reasoning for Multi-modal Large Language Model
Byungwoo Jeon, Yoonwoo Jeong, Hyunseok Lee +2
Despite recent advancements in Multi-modal Large Language Models (MLLMs) on diverse understanding tasks, these models struggle to solve problems which require extensive multi-step…
RoDyGS: Robust Dynamic Gaussian Splatting for Casual Videos
Junmyeong Lee, Hoseung Choi, Yoonwoo Jeong +1
4D reconstruction from casually captured monocular videos is challenging due to inherent ambiguity in reconstructing dynamic 3D geometry. To address this challenge, we introduce Ro…
DextER: Language-driven Dexterous Grasp Generation with Embodied Reasoning
Junha Lee, Eunha Park, Minsu Cho
Language-driven dexterous grasp generation requires the models to understand task semantics, 3D geometry, and complex hand-object interactions. While vision-language models have be…
Affostruction: 3D Affordance Grounding with Generative Reconstruction
Chunghyun Park, Seunghyeon Lee, Minsu Cho
This paper addresses the problem of affordance grounding from RGBD images of an object, which aims to localize surface regions corresponding to a text query that describes an actio…
MV-SAM: Multi-view Promptable Segmentation using Pointmap Guidance
Yoonwoo Jeong, Cheng Sun, Yu-Chiang Frank Wang +2
Promptable segmentation has emerged as a powerful paradigm in computer vision, enabling users to guide models in parsing complex scenes with prompts such as clicks, boxes, or textu…