7 papers · 1 filter
JetViT: Efficient High-Resolution Vision Transformer with Post-Training Attention Search
Dongyun Zou, Zhuoyang Zhang, Junyu Chen +8
We introduce JetViT, a novel family of hybrid-architecture Vision Transformer (ViT) models that match the accuracy of state-of-the-art full-attention vision foundation models while…
Locality-aware Parallel Decoding for Efficient Autoregressive Image Generation
Zhuoyang Zhang, Luke J. Huang, Chengyue Wu +4
We present Locality-aware Parallel Decoding (LPD) to accelerate autoregressive image generation. Traditional autoregressive image generation relies on next-patch prediction, a memo…
DC-AR: Efficient Masked Autoregressive Image Generation with Deep Compression Hybrid Tokenizer
Yecheng Wu, Junyu Chen, Zhuoyang Zhang +7
We introduce DC-AR, a novel masked autoregressive (AR) text-to-image generation framework that delivers superior image generation quality with exceptional computational efficiency.…
Deep Compression Autoencoder for Efficient High-Resolution Diffusion Models
Junyu Chen, Han Cai, Junsong Chen +6
We present Deep Compression Autoencoder (DC-AE), a new family of autoencoder models for accelerating high-resolution diffusion models. Existing autoencoder models have demonstrated…
CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models
Qingqing Zhao, Yao Lu, Moo Jin Kim +12
Vision-language-action models (VLAs) have shown potential in leveraging pretrained vision-language models and diverse robot demonstrations for learning generalizable sensorimotor c…
SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformers
Enze Xie, Junsong Chen, Junyu Chen +8
We introduce Sana, a text-to-image framework that can efficiently generate images up to 40964096 resolution. Sana can synthesize high-resolution, high-quality images with s…