3 papers
cs.CV2025
Denoising Vision Transformer Autoencoder with Spectral Self-Regularization
Xunzhi Xiang, Xingye Tian, Guiyu Zhang +5
Variational autoencoders (VAEs) typically encode images into a compact latent space, reducing computational cost but introducing an optimization dilemma: a higher-dimensional laten…
cs.CV2025
Macro-from-Micro Planning for High-Quality and Parallelized Autoregressive Long Video Generation
Xunzhi Xiang, Yabo Chen, Guiyu Zhang +10
Current autoregressive diffusion models excel at video generation but are generally limited to short temporal durations. Our theoretical analysis indicates that the autoregressive…
cs.CV2025
Make It Efficient: Dynamic Sparse Attention for Autoregressive Image Generation
Xunzhi Xiang, Qi Fan
Autoregressive conditional image generation models have emerged as a dominant paradigm in text-to-image synthesis. These methods typically convert images into one-dimensional token…