activity
20222024
most citedDiffusion-RWKV: Scaling RWKV-Like Architectures for Diffusion Models

5 citations · 13 across the 8 of their papers we have counts for

collaborators

8 papers

cs.CV2024

Scaling Diffusion Transformers to 16 Billion Parameters

Zhengcong Fei, Mingyuan Fan, Changqian Yu +2

In this paper, we present DiT-MoE, a sparse version of the diffusion Transformer, that is scalable and competitive with dense networks while exhibiting highly optimized inference.…

cs.SD2024

Music Consistency Models

Zhengcong Fei, Mingyuan Fan, Junshi Huang

Consistency models have exhibited remarkable capabilities in facilitating efficient image/video generation, enabling synthesis with minimal sampling steps. It has proven to be adva…

cs.CV20245 cited

Diffusion-RWKV: Scaling RWKV-Like Architectures for Diffusion Models

Zhengcong Fei, Mingyuan Fan, Changqian Yu +2

Transformers have catalyzed advancements in computer vision and natural language processing (NLP) fields. However, substantial computational complexity poses limitations for their…

cs.CV20242 cited

Scalable Diffusion Models with State Space Backbone

Zhengcong Fei, Mingyuan Fan, Changqian Yu +1

This paper presents a new exploration into a category of diffusion models built upon state space architecture. We endeavor to train diffusion models for image data, wherein the tra…

cs.CV20233 cited

Prefix-diffusion: A Lightweight Diffusion Model for Diverse Image Captioning

Guisheng Liu, Yi Li, Zhengcong Fei +3

While impressive performance has been achieved in image captioning, the limited diversity of the generated captions and the large parameter scale remain major barriers to the real-…

cs.CV20232 cited

DiT: Efficient Vision Transformers with Dynamic Token Routing

Yuchen Ma, Zhengcong Fei, Junshi Huang

Recently, the tokens of images share the same static data flow in many dense networks. However, challenges arise from the variance among the objects in images, such as large variat…