8 papers
Alignment-Guided Score Matching for Text-to-Image Alignment in Diffusion Models
Jaa-Yeon Lee, Yeobin Hong, Taesung Kwon +1
Diffusion models generate highly realistic images but often struggle with precise text-image alignment. While recent post-training methods improve alignment using external rewards…
Accelerating Video Inverse Problem Solvers with Autoregressive Diffusion Models
Taesung Kwon, Jonghyun Park, Hyungjin Chung +1
Diffusion models provide powerful priors for zero-shot video inverse problems, but their real-time deployment is hindered by two inefficiencies: high initial latency caused by holi…
Geometric 4D Stitching for Grounded 4D Generation
Sunwoo Park, Taesung Kwon, Jong Chul Ye
Recent 4D generation methods complete scene-level missing information using generative models and reconstruct the scene into radiance-based representations. However, these pipeline…
Reviving ConvNeXt for Efficient Convolutional Diffusion Models
Taesung Kwon, Lorenzo Bianchi, Lennart Wittke +5
Recent diffusion models increasingly favor Transformer backbones, motivated by the remarkable scalability of fully attentional architectures. Yet the locality bias, parameter effic…
Zero4D: Training-Free 4D Video Generation From Single Video Using Off-the-Shelf Video Diffusion
Jangho Park, Taesung Kwon, Jong Chul Ye
Multi-view or 4D video generation has emerged as a significant research topic. Nonetheless, recent approaches to 4D generation still struggle with fundamental limitations, as they…
VISION-XL: High Definition Video Inverse Problem Solver using Latent Image Diffusion Models
Taesung Kwon, Jong Chul Ye
In this paper, we propose a novel framework for solving high-definition video inverse problems using latent image diffusion models. Building on recent advancements in spatio-tempor…