7 papers · 1 filter
Group3D: MLLM-Driven Semantic Grouping for Open-Vocabulary 3D Object Detection
Youbin Kim, Jinho Park, Hogun Park +1
Open-vocabulary 3D object detection aims to localize and recognize objects beyond a fixed training taxonomy. In multi-view RGB settings, recent approaches often decouple geometry-b…
3DreamBooth: High-Fidelity 3D Subject-Driven Video Generation Model
Hyun-kyu Ko, Jihyeon Park, Younghyun Kim +2
Creating dynamic, view-consistent videos of customized subjects is highly sought after for a wide range of emerging applications, including immersive VR/AR, virtual production, and…
OpenMonoGS-SLAM: Monocular Gaussian Splatting SLAM with Open-set Semantics
Jisang Yoo, Gyeongjin Kang, Hyun-kyu Ko +2
Simultaneous Localization and Mapping (SLAM) is a foundational component in robotics, AR/VR, and autonomous systems. With the rising focus on spatial AI in recent years, combining…
Gather-Scatter Mamba: Accelerating Propagation with Efficient State Space Model
Hyun-kyu Ko, Youbin Kim, Jihyeon Park +5
State Space Models (SSMs)-most notably RNNs-have historically played a central role in sequential modeling. Although attention mechanisms such as Transformers have since dominated…
DiffuseSlide: Training-Free High Frame Rate Video Generation Diffusion
Geunmin Hwang, Hyun-kyu Ko, Younghyun Kim +2
Recent advancements in diffusion models have revolutionized video generation, enabling the creation of high-quality, temporally consistent videos. However, generating high frame-ra…
CompMarkGS: Robust Watermarking for Compressed 3D Gaussian Splatting
Sumin In, Youngdong Jang, Utae Jeong +4
As 3D Gaussian Splatting (3DGS) is increasingly adopted in various academic and commercial applications due to its high-quality and real-time rendering capabilities, the need for c…