From the 1 of 9 linked papers with an AI index.
9 papers
SplatReasoner: Enhancing Embodied Reasoning and Grounding by Novel View Synthesis
Kim Yu-Ji, Dahye Lee, Kim Jun-Seong +6
SplatReasoner integrates query‑conditioned novel view synthesis via 3D Gaussian splatting into vision‑language models, enabling better embodied reasoning and 3D grounding by genera…
DVSM: Decoder-only View Synthesis Model Done Right
Cheng Sun, Jaesung Choe, Min-Hung Chen +2
Recent Large View Synthesis Models (LVSMs) advocate an encoder-decoder architecture that separates reconstruction and rendering into distinct networks. We re-examine this design. T…
MV-SAM: Multi-view Promptable Segmentation using Pointmap Guidance
Yoonwoo Jeong, Cheng Sun, Yu-Chiang Frank Wang +2
Promptable segmentation has emerged as a powerful paradigm in computer vision, enabling users to guide models in parsing complex scenes with prompts such as clicks, boxes, or textu…
OpenVoxel: Training-Free Grouping and Captioning Voxels for Open-Vocabulary 3D Scene Understanding
Sheng-Yu Huang, Jaesung Choe, Yu-Chiang Frank Wang +1
We propose OpenVoxel, a training-free algorithm for grouping and captioning sparse voxels for the open-vocabulary 3D scene understanding tasks. Given the sparse voxel rasterization…
Quantile Rendering: Efficiently Embedding High-dimensional Feature on 3D Gaussian Splatting
Yoonwoo Jeong, Cheng Sun, Frank Wang +2
Recent advancements in computer vision have successfully extended Open-vocabulary segmentation (OVS) to the 3D domain by leveraging 3D Gaussian Splatting (3D-GS). Despite this prog…
SVRecon: Sparse Voxel Rasterization for Surface Reconstruction
Seunghun Oh, Jaesung Choe, Dongjae Lee +4
We extend the recently proposed sparse voxel rasterization paradigm to the task of high-fidelity surface reconstruction by integrating Signed Distance Function (SDF), named SVRecon…