16 papers
TanGO: Training-Free 3D Editing via Tangent-Space Guidance and Optimization
Siwoo Lim, Sunjae Yoon, Gwanhyeong Koo +2
While recent flow-matching 3D generative models (e.g., VecSet) adopt structured representations, their tokens share global context, causing conventional training-free editing to su…
InSpace: Structure-Aware 3D Indoor Scene Generation from a Single 360° Image
Gwanhyeong Koo, Hyunsu Kim, Youngji Kim +5
Recent advances in single image-to-3D generation have enabled high-quality asset synthesis, yet extending these capabilities to indoor scene generation remains challenging. Existin…
GADA: Geometry-Aware Deformable Aggregation for Image-Based Gaussian Splatting
Siwoo Lim, Sunjae Yoon, Gwanhyeong Koo +1
Gaussian Splatting has achieved significant improvements by incorporating warping-based techniques. However, such methods suffer from pixel-level inaccuracies due to uncertain geom…
Query-based Cross-Modal Projector Bolstering Mamba Multimodal LLM
SooHwan Eom, Jay Shim, Gwanhyeong Koo +4
The Transformer's quadratic complexity with input length imposes an unsustainable computational load on large language models (LLMs). In contrast, the Selective Scan Structured Sta…
PDCR: Perception-Decomposed Confidence Reward for Vision-Language Reasoning
Hee Suk Yoon, Eunseop Yoon, Ji Woo Hong +6
Reinforcement Learning with Verifiable Rewards (RLVR) traditionally relies on a sparse, outcome-based signal. Recent work shows that providing a fine-grained, model-intrinsic signa…
High-Fidelity Text-to-Image Generation from Pre-Trained Vision-Language Models via Distribution-Conditioned Diffusion Decoding
Ji Woo Hong, Hee Suk Yoon, Gwanhyeong Koo +5
Recent large-scale vision-language models (VLMs) have shown remarkable text-to-image generation capabilities, yet their visual fidelity remains constrained by the discrete image to…