activity
20242026
collaborators

8 papers

cs.CV2026

Boosting Resolution Generalization of Diffusion Transformers with Randomized Positional Encodings

Liang Hou, Cong Liu, Mingwu Zheng +4

Resolution generalization in image generation tasks enables the production of higher-resolution images with lower training resolution overhead. However, a key obstacle for diffusio…

cs.CV2025

Improving Video Generation with Human Feedback

Jie Liu, Gongye Liu, Jiajun Liang +14

Video generation has achieved significant advances through rectified flow techniques, but issues like unsmooth motion and misalignment between videos and prompts persist. In this w…

cs.CV2025

Koala-36M: A Large-scale Video Dataset Improving Consistency between Fine-grained Conditions and Video Content

Qiuheng Wang, Yukai Shi, Jiarong Ou +10

With the continuous progress of visual generation technologies, the scale of video datasets has grown exponentially. The quality of these datasets plays a pivotal role in the perfo…

cs.CV2025

DiffMoE: Dynamic Token Selection for Scalable Diffusion Transformers

Minglei Shi, Ziyang Yuan, Haotian Yang +10

Diffusion models have demonstrated remarkable success in various image generation tasks, but their performance is often limited by the uniform processing of inputs across varying c…

cs.CV2024

Towards Precise Scaling Laws for Video Diffusion Transformers

Yuanyang Yin, Yaqi Zhao, Mingwu Zheng +11

Achieving optimal performance of video diffusion transformers within given data and compute budget is crucial due to their high training costs. This necessitates precisely determin…

cs.CV2024

ImFace++: A Sophisticated Nonlinear 3D Morphable Face Model with Implicit Neural Representations

Mingwu Zheng, Haiyu Zhang, Hongyu Yang +2

Accurate representations of 3D faces are of paramount importance in various computer vision and graphics applications. However, the challenges persist due to the limitations impose…