5 papers
Nested AutoRegressive Models
Hongyu Wu, Xuhui Fan, Zhangkai Wu +1
AutoRegressive (AR) models have demonstrated competitive performance in image generation, achieving results comparable to those of diffusion models. However, their token-by-token i…
VALA: Learning Latent Anchors for Training-Free and Temporally Consistent
Zhangkai Wu, Xuhui Fan, Zhongyuan Xie +2
Recent advances in training-free video editing have enabled lightweight and precise cross-frame generation by leveraging pre-trained text-to-image diffusion models. However, existi…
FAME: Fairness-aware Attention-modulated Video Editing
Zhangkai Wu, Xuhui Fan, Zhongyuan Xie +3
Training-free video editing (VE) models tend to fall back on gender stereotypes when rendering profession-related prompts. We propose \textbf{FAME} for \textit{Fairness-aware Atten…
SCoT: Unifying Consistency Models and Rectified Flows via Straight-Consistent Trajectories
Zhangkai Wu, Xuhui Fan, Hongyu Wu +1
Pre-trained diffusion models are commonly used to generate clean data (e.g., images) from random noises, effectively forming pairs of noises and corresponding clean images. Distill…
Marked Temporal Bayesian Flow Point Processes
Hui Chen, Xuhui Fan, Hengyu Liu +1
Marked event data captures events by recording their continuous-valued occurrence timestamps along with their corresponding discrete-valued types. They have appeared in various rea…