11 papers
Multi-view Pyramid Transformer: Look Coarser to See Broader
Gyeongjin Kang, Seungkwon Yang, Seungtae Nam +3
We propose Multi-view Pyramid Transformer (MVP), a scalable multi-view transformer architecture that directly reconstructs large 3D scenes from tens to hundreds of images in a sing…
VGenST-Bench: A Benchmark for Spatio-Temporal Reasoning via Active Video Synthesis
Jinho Park, Youbin Kim, Hogun Park +1
Spatio-temporal reasoning is a core capability for Multimodal Large Language Models (MLLMs) operating in the real world. As such, evaluating it precisely has become an essential ch…
2Xplat: Decoupling Geometry and Appearance Modeling for Feed-Forward 3D Gaussian Splatting
Hwasik Jeong, Seungryong Lee, Gyeongjin Kang +4
Pose-free feed-forward 3D Gaussian Splatting (3DGS) has opened a new frontier for rapid 3D modeling, enabling high-quality Gaussian representations to be generated from uncalibrate…
Group3D: MLLM-Driven Semantic Grouping for Open-Vocabulary 3D Object Detection
Youbin Kim, Jinho Park, Hogun Park +1
Open-vocabulary 3D object detection aims to localize and recognize objects beyond a fixed training taxonomy. In multi-view RGB settings, recent approaches often decouple geometry-b…
3DreamBooth: High-Fidelity 3D Subject-Driven Video Generation Model
Hyun-kyu Ko, Jihyeon Park, Younghyun Kim +2
Creating dynamic, view-consistent videos of customized subjects is highly sought after for a wide range of emerging applications, including immersive VR/AR, virtual production, and…
OpenMonoGS-SLAM: Monocular Gaussian Splatting SLAM with Open-set Semantics
Jisang Yoo, Gyeongjin Kang, Hyun-kyu Ko +2
Simultaneous Localization and Mapping (SLAM) is a foundational component in robotics, AR/VR, and autonomous systems. With the rising focus on spatial AI in recent years, combining…