4 papers · 1 filter
JetViT: Efficient High-Resolution Vision Transformer with Post-Training Attention Search
Dongyun Zou, Zhuoyang Zhang, Junyu Chen +8
We introduce JetViT, a novel family of hybrid-architecture Vision Transformer (ViT) models that match the accuracy of state-of-the-art full-attention vision foundation models while…
UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation
Jiayun Wang, Yu Wang, Weijie Gan +2
We introduce spatially grounded contextual image generation, a controllable image generation task that reframes the conditioning paradigm. Instead of supplying a reference image an…
DiP-GO: A Diffusion Pruner via Few-step Gradient Optimization
Haowei Zhu, Dehua Tang, Ji Liu +12
Diffusion models have achieved remarkable progress in the field of image generation due to their outstanding capabilities. However, these models require substantial computing resou…
UPDP: A Unified Progressive Depth Pruner for CNN and Vision Transformer
Ji Liu, Dehua Tang, Yuanxian Huang +9
Traditional channel-wise pruning methods by reducing network channels struggle to effectively prune efficient CNN models with depth-wise convolutional layers and certain efficient…