4 papers · 1 filter
Nested AutoRegressive Models
Hongyu Wu, Xuhui Fan, Zhangkai Wu +1
AutoRegressive (AR) models have demonstrated competitive performance in image generation, achieving results comparable to those of diffusion models. However, their token-by-token i…
VALA: Learning Latent Anchors for Training-Free and Temporally Consistent
Zhangkai Wu, Xuhui Fan, Zhongyuan Xie +2
Recent advances in training-free video editing have enabled lightweight and precise cross-frame generation by leveraging pre-trained text-to-image diffusion models. However, existi…
FAME: Fairness-aware Attention-modulated Video Editing
Zhangkai Wu, Xuhui Fan, Zhongyuan Xie +3
Training-free video editing (VE) models tend to fall back on gender stereotypes when rendering profession-related prompts. We propose \textbf{FAME} for \textit{Fairness-aware Atten…
SCoT: Unifying Consistency Models and Rectified Flows via Straight-Consistent Trajectories
Zhangkai Wu, Xuhui Fan, Hongyu Wu +1
Pre-trained diffusion models are commonly used to generate clean data (e.g., images) from random noises, effectively forming pairs of noises and corresponding clean images. Distill…