16 papers
Distributional Matching for Vector Quantization: A Unified Theoretical and Empirical Framework
Xianghong Fang, Litao Guo, Hengchao Chen +8
The effectiveness of modern visual representation learning and autoregressive models critically depends on vector quantization (VQ), which discretizes continuous feature representa…
RotateAttention: RoPE-Aware Rotation and Range Rectification for INT4 Quantized Attention in Video Generation
Yaofu Liu, Wanli Lan, Jinxi Li +2
In , the attention mechanism remains a primary computational bottleneck due to its…
GKDT: General Keypoint Detection Transformer
Changsheng Lu, Yuxin Chen, Haokun Gui +5
With the emergence of various pre-trained vision and language models, computer vision is shifting from narrow-domain to open-domain recognition. The construction of a more powerful…
ScalingAR: Scaling Confidence for Autoregressive Image Generation
Harold Haodong Chen, Xianfeng Wu, Wen-Jie Shu +4
Test-time strategies have shown remarkable success in improving large language models, but their application to next-token prediction (NTP) autoregressive (AR) image generation rem…
AlignVid: Training-Free Attention Scaling for Semantic Fidelity in Text-Guided Image-to-Video Generation
Yexin Liu, Wen-Jie Shu, Zile Huang +6
Text-guided image-to-video generation has made substantial progress, yet it still struggles to execute text-specified edits that require substantial changes to a reference image (\…
AHPA: Adaptive Hierarchical Prior Alignment for Diffusion Transformers
Ruibin Min, Yexin Liu, Aimin Pan +5
Representation alignment has recently emerged as an effective paradigm for accelerating Diffusion Transformer training. Despite their success, existing alignment methods typically…