1 citations · 1 across the 2 of their papers we have counts for
3 papers
cs.CV2025
Playing with Transformer at 30+ FPS via Next-Frame Diffusion
Xinle Cheng, Tianyu He, Jiayi Xu +3
Autoregressive video models offer distinct advantages over bidirectional diffusion models in creating interactive video content and supporting streaming applications with arbitrary…
cs.CV2025★ 1 cited
CAT Pruning: Cluster-Aware Token Pruning For Text-to-Image Diffusion Models
Xinle Cheng, Zhuoming Chen, Zhihao Jia
Diffusion models have revolutionized generative tasks, especially in the domain of text-to-image synthesis; however, their iterative denoising process demands substantial computati…
cs.CV2024
VidTok: A Versatile and Open-Source Video Tokenizer
Anni Tang, Tianyu He, Junliang Guo +3
Encoding video content into compact latent tokens has become a fundamental step in video generation and understanding, driven by the need to address the inherent redundancy in pixe…