5 papers
A Renaissance of Explicit Motion Information Mining from Transformers for Action Recognition
Peiqin Zhuang, Lei Bai, Yichao Wu +4
Recently, action recognition has been dominated by transformer-based methods, thanks to their spatiotemporal contextual aggregation capacities. However, despite the significant pro…
Understand Before You Generate: Self-Guided Training for Autoregressive Image Generation
Xiaoyu Yue, Zidong Wang, Yuqing Wang +5
Recent studies have demonstrated the importance of high-quality visual representations in image generation and have highlighted the limitations of generative models in image unders…
InfGen: A Resolution-Agnostic Paradigm for Scalable Image Synthesis
Tao Han, Wanghan Xu, Junchao Gong +4
Arbitrary resolution image generation provides a consistent visual experience across devices, having extensive applications for producers and consumers. Current diffusion models in…
Exploring Representation-Aligned Latent Space for Better Generation
Wanghan Xu, Xiaoyu Yue, Zidong Wang +6
Generative models serve as powerful tools for modeling the real world, with mainstream diffusion models, particularly those based on the latent diffusion model paradigm, achieving…
Diffusion Models Need Visual Priors for Image Generation
Xiaoyu Yue, Zidong Wang, Zeyu Lu +5
Conventional class-guided diffusion models generally succeed in generating images with correct semantic content, but often struggle with texture details. This limitation stems from…