collaborators
Showing cs.CVShow all

6 papers · 1 filter

cs.CV2026

CascadeOcc: Rethinking 3D Occupancy World Models with Cascaded VQ Representations

Kyumin Hwang, Wonhyeok Choi, Jaeyeul Kim +3

This letter proposes CascadeOcc, a novel occupancy world model that prioritizes intrinsic structural hierarchy over extrinsic auxiliary modalities for autonomous driving. Occupancy…

cs.CV2025

Scale-invariant and View-relational Representation Learning for Full Surround Monocular Depth

Kyumin Hwang, Wonhyeok Choi, Kiljoon Han +5

Recent foundation models demonstrate strong generalization capabilities in monocular depth estimation. However, directly applying these models to Full Surround Monocular Depth Esti…

cs.CV2025

Infinite-Story: A Training-Free Consistent Text-to-Image Generation

Jihun Park, Kyoungmin Lee, Jongmin Gim +7

We present Infinite-Story, a training-free framework for consistent text-to-image (T2I) generation tailored for multi-prompt storytelling scenarios. Built upon a scale-wise autoreg…

cs.CV2025

A Training-Free Style-Personalization via SVD-Based Feature Decomposition

Kyoungmin Lee, Jihun Park, Jongmin Gim +4

We present a training-free framework for style-personalized image generation that operates during inference using a scale-wise autoregressive model. Our method generates a stylized…

cs.CV2025

Bridging Geometric and Semantic Foundation Models for Generalized Monocular Depth Estimation

Sanggyun Ma, Wonjoon Choi, Jihun Park +4

We present Bridging Geometric and Semantic (BriGeS), an effective method that fuses geometric and semantic information within foundation models to enhance Monocular Depth Estimatio…

cs.CV2025

A Training-Free Style-aligned Image Generation with Scale-wise Autoregressive Model

Jihun Park, Jongmin Gim, Kyoungmin Lee +5

We present a training-free style-aligned image generation method that leverages a scale-wise autoregressive model. While large-scale text-to-image (T2I) models, particularly diffus…