103 citations · 214 across the 13 of their papers we have counts for
21 papers
AGFT: Alignment-Guided Fine-Tuning for Zero-Shot Adversarial Robustness of Vision-Language Models
Yubo Cui, Xianchao Guan, Zijun Xiong +1
Pre-trained vision-language models (VLMs) exhibit strong zero-shot generalization but remain vulnerable to adversarial perturbations. Existing classification-guided adversarial fin…
Rényi Entropy: A New Token Pruning Metric for Vision Transformers
Wei-Yuan Su, Ruijie Zhang, Zheng Zhang
Vision Transformers (ViTs) achieve state-of-the-art performance but suffer from the complexity of self-attention, making inference costly for high-resolution inputs. To ad…
DETR Doesn't Need Multi-Scale or Locality Design
Yutong Lin, Yuhui Yuan, Zheng Zhang +3
This paper presents an improved DETR detector that maintains a "plain" nature: using a single-scale feature map and global cross-attention calculations without specific locality co…
Exploring Discrete Diffusion Models for Image Captioning
Zixin Zhu, Yixuan Wei, Jianfeng Wang +7
The image captioning task is typically realized by an auto-regressive method that decodes the text tokens one by one. We present a diffusion-based captioning model, dubbed the name…
Could Giant Pretrained Image Models Extract Universal Representations?
Yutong Lin, Ze Liu, Zheng Zhang +4
Frozen pretrained models have become a viable alternative to the pretraining-then-finetuning paradigm for transfer learning. However, with frozen models there are relatively few pa…
Expediting Large-Scale Vision Transformer for Dense Prediction without Fine-tuning
Weicong Liang, Yuhui Yuan, Henghui Ding +6
Vision transformers have recently achieved competitive results across various vision tasks but still suffer from heavy computation costs when processing a large number of tokens. M…