8 papers · 1 filter
CascadeOcc: Rethinking 3D Occupancy World Models with Cascaded VQ Representations
Kyumin Hwang, Wonhyeok Choi, Jaeyeul Kim +3
This letter proposes CascadeOcc, a novel occupancy world model that prioritizes intrinsic structural hierarchy over extrinsic auxiliary modalities for autonomous driving. Occupancy…
Scale-invariant and View-relational Representation Learning for Full Surround Monocular Depth
Kyumin Hwang, Wonhyeok Choi, Kiljoon Han +5
Recent foundation models demonstrate strong generalization capabilities in monocular depth estimation. However, directly applying these models to Full Surround Monocular Depth Esti…
A Training-Free Style-Personalization via SVD-Based Feature Decomposition
Kyoungmin Lee, Jihun Park, Jongmin Gim +4
We present a training-free framework for style-personalized image generation that operates during inference using a scale-wise autoregressive model. Our method generates a stylized…
Infinite-Story: A Training-Free Consistent Text-to-Image Generation
Jihun Park, Kyoungmin Lee, Jongmin Gim +7
We present Infinite-Story, a training-free framework for consistent text-to-image (T2I) generation tailored for multi-prompt storytelling scenarios. Built upon a scale-wise autoreg…
Temporal Grounding as a Learning Signal for Referring Video Object Segmentation
Seunghun Lee, Jiwan Seo, Jeonghoon Kim +9
Referring Video Object Segmentation (RVOS) aims to segment and track objects in videos based on natural language expressions, requiring precise alignment between visual content and…
Intrinsic Image Decomposition for Robust Self-supervised Monocular Depth Estimation on Reflective Surfaces
Wonhyeok Choi, Kyumin Hwang, Minwoo Choi +4
Self-supervised monocular depth estimation (SSMDE) has gained attention in the field of deep learning as it estimates depth without requiring ground truth depth maps. This approach…