6 papers
Diffusing in the Right Space: A Systematic Study of Latent Diffusability
Tianxiong Zhong, Xingye Tian, Xuebo Wang +2
Latent diffusion models leverage visual tokenizers to compress images into latent spaces for efficient generative modeling. However, better reconstruction quality of a tokenizer do…
CamPilot: Improving Camera Control in Video Diffusion Model with Efficient Camera Reward Feedback
Wenhang Ge, Guibao Shen, Jiawei Feng +5
Recent advances in camera-controlled video diffusion models have significantly improved video-camera alignment. However, the camera controllability still remains limited. In this w…
Decoupling Complexity from Scale in Latent Diffusion Model
Tianxiong Zhong, Xingye Tian, Xuebo Wang +3
Existing latent diffusion models typically couple scale with content complexity, using more latent tokens to represent higher-resolution images or higher-frame rate videos. However…
Denoising Vision Transformer Autoencoder with Spectral Self-Regularization
Xunzhi Xiang, Xingye Tian, Guiyu Zhang +5
Variational autoencoders (VAEs) typically encode images into a compact latent space, reducing computational cost but introducing an optimization dilemma: a higher-dimensional laten…
VFRTok: Variable Frame Rates Video Tokenizer with Duration-Proportional Information Assumption
Tianxiong Zhong, Xingye Tian, Boyuan Jiang +4
Modern video generation frameworks based on Latent Diffusion Models suffer from inefficiencies in tokenization due to the Frame-Proportional Information Assumption. Existing tokeni…
VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing
Jiahao Hu, Tianxiong Zhong, Xuebo Wang +5
Diffusion-based image editing models have made remarkable progress in recent years. However, achieving high-quality video editing remains a significant challenge. One major hurdle…