collaborators

5 papers

cs.CV2025

A Renaissance of Explicit Motion Information Mining from Transformers for Action Recognition

Peiqin Zhuang, Lei Bai, Yichao Wu +4

Recently, action recognition has been dominated by transformer-based methods, thanks to their spatiotemporal contextual aggregation capacities. However, despite the significant pro…

cs.CV2025

Understand Before You Generate: Self-Guided Training for Autoregressive Image Generation

Xiaoyu Yue, Zidong Wang, Yuqing Wang +5

Recent studies have demonstrated the importance of high-quality visual representations in image generation and have highlighted the limitations of generative models in image unders…

cs.CV2025

InfGen: A Resolution-Agnostic Paradigm for Scalable Image Synthesis

Tao Han, Wanghan Xu, Junchao Gong +4

Arbitrary resolution image generation provides a consistent visual experience across devices, having extensive applications for producers and consumers. Current diffusion models in…

cs.LG2025

Exploring Representation-Aligned Latent Space for Better Generation

Wanghan Xu, Xiaoyu Yue, Zidong Wang +6

Generative models serve as powerful tools for modeling the real world, with mainstream diffusion models, particularly those based on the latent diffusion model paradigm, achieving…

cs.CV2024

Diffusion Models Need Visual Priors for Image Generation

Xiaoyu Yue, Zidong Wang, Zeyu Lu +5

Conventional class-guided diffusion models generally succeed in generating images with correct semantic content, but often struggle with texture details. This limitation stems from…