10 papers
Transferability Between Understanding and Generation in Unified Multimodal Models
Jiwon Kang, Heeji Yoon, Jaewoo Jung +5
Unified Multimodal Models (UMMs) integrate image understanding and generation within a single architecture, yet how the two tasks interact remains understudied. We investigate $\bo…
Probing and Leveraging Video Diffusion Transformer Features for Robust Point Tracking
Soowon Son, Honggyu An, Jisu Nam +7
Despite achieving strong results on standard benchmarks, current point tracking methods rely on feature backbones that are rarely designed with the temporal coherence needed for ro…
GeoFace: Consistent Multi-View Face Generation with Geometry-Constrained Diffusion
Yeji Choi, Jinhyeok Choi, Jaewon Min +3
We present GeoFace, a geometry-constrained multi-view diffusion framework for consistent face generation from a single input. % While recent multi-view diffusion models achieve pho…
Geometry-Aware Representation Denoising for Robust Multi-view 3D Reconstruction
Jin Hyeon Kim, Jaeeun Lee, Claire Kim +8
Multi-view 3D reconstruction has achieved remarkable progress with the advent of feed-forward 3D reconstruction models. However, these models are typically trained and evaluated un…
DA-Flow: Degradation-Aware Optical Flow Estimation with Diffusion Models
Jaewon Min, Jaeeun Lee, Yeji Choi +7
Optical flow models trained on high-quality data often degrade severely when confronted with real-world corruptions such as blur, noise, and compression artifacts. To overcome this…
Unified Diffusion Transformer for High-fidelity Text-Aware Image Restoration
Jin Hyeon Kim, Paul Hyunbin Cho, Claire Kim +5
Text-Aware Image Restoration (TAIR) aims to recover high-quality images from low-quality inputs containing degraded textual content. While diffusion models provide strong generativ…