collaborators

7 papers

cs.CV2025

Zero-Shot Video Translation and Editing with Frame Spatial-Temporal Correspondence

Shuai Yang, Junxin Lin, Yifan Zhou +2

The remarkable success in text-to-image diffusion models has motivated extensive investigation of their potential for video applications. Zero-shot techniques aim to adapt image di…

cs.CV2025

WorldMem: Long-term Consistent World Simulation with Memory

Zeqi Xiao, Yushi Lan, Yifan Zhou +4

World simulation has gained increasing popularity due to its ability to model virtual environments and predict the consequences of actions. However, the limited temporal context wi…

cs.CV2025

FreeFlux: Understanding and Exploiting Layer-Specific Roles in RoPE-Based MMDiT for Versatile Image Editing

Tianyi Wei, Yifan Zhou, Dongdong Chen +1

The integration of Rotary Position Embedding (RoPE) in Multimodal Diffusion Transformer (MMDiT) has significantly enhanced text-to-image generation quality. However, the fundamenta…

cs.CV2025

Alias-Free Latent Diffusion Models: Improving Fractional Shift Equivariance of Diffusion Latent Space

Yifan Zhou, Zeqi Xiao, Shuai Yang +1

Latent Diffusion Models (LDMs) are known to have an unstable generation process, where even small perturbations or shifts in the input noise can lead to significantly different out…

cs.CV2024

Trajectory Attention for Fine-grained Video Motion Control

Zeqi Xiao, Wenqi Ouyang, Yifan Zhou +4

Recent advancements in video generation have been greatly driven by video diffusion models, with camera motion control emerging as a crucial challenge in creating view-customized v…

cs.CV2024

Enhancing MMDiT-Based Text-to-Image Models for Similar Subject Generation

Tianyi Wei, Dongdong Chen, Yifan Zhou +1

Representing the cutting-edge technique of text-to-image models, the latest Multimodal Diffusion Transformer (MMDiT) largely mitigates many generation issues existing in previous m…