4 papers
Aligning Latent Spaces with Flow Priors
Yizhuo Li, Yuying Ge, Yixiao Ge +2
This paper presents a novel framework for aligning learnable latent spaces to arbitrary target distributions by leveraging flow-based generative models as priors. Our method first…
DiCoDe: Diffusion-Compressed Deep Tokens for Autoregressive Video Generation with Language Models
Yizhuo Li, Yuying Ge, Yixiao Ge +2
Videos are inherently temporal sequences by their very nature. In this work, we explore the potential of modeling videos in a chronological and scalable manner with autoregressive…
StyleAdapter: A Unified Stylized Image Generation Model
Zhouxia Wang, Xintao Wang, Liangbin Xie +4
This work focuses on generating high-quality images with specific style of reference images and content of provided textual descriptions. Current leading algorithms, i.e., DreamBoo…
Analysis and Benchmarking of Extending Blind Face Image Restoration to Videos
Zhouxia Wang, Jiawei Zhang, Xintao Wang +4
Recent progress in blind face restoration has resulted in producing high-quality restored results for static images. However, efforts to extend these advancements to video scenario…