5 papers
RayRoPE: Projective Ray Positional Encoding for Multi-view Attention
Yu Wu, Minsik Jeon, Jen-Hao Rick Chang +2
We study positional encodings for multi-view transformers that process tokens from a set of posed input images, and seek a mechanism that encodes patches uniquely, allows SE(3)-inv…
Novel View Synthesis as Video Completion
Qi Wu, Khiem Vuong, Minsik Jeon +2
We tackle the problem of sparse novel view synthesis (NVS) using video diffusion models; given () multi-view images of a scene and their camera poses, we predict the…
Flow3r: Factored Flow Prediction for Scalable Visual Geometry Learning
Zhongxiao Cong, Qitao Zhao, Minsik Jeon +1
Current feed-forward 3D/4D reconstruction systems rely on dense geometry and pose supervision -- expensive to obtain at scale and particularly scarce for dynamic real-world scenes.…
E2-BKI: Evidential Ellipsoidal Bayesian Kernel Inference for Uncertainty-aware Gaussian Semantic Mapping
Junyoung Kim, Minsik Jeon, Jihong Min +2
Semantic mapping aims to construct a 3D semantic representation of the environment, providing essential knowledge for robots operating in complex outdoor settings. While Bayesian K…
OW-Rep: Open World Object Detection with Instance Representation Learning
Sunoh Lee, Minsik Jeon, Jihong Min +1
Open World Object Detection(OWOD) addresses realistic scenarios where unseen object classes emerge, enabling detectors trained on known classes to detect unknown objects and increm…