activity
20182026
most citedAn Empirical Study of Spatial Attention Mechanisms in Deep Networks

103 citations · 214 across the 13 of their papers we have counts for

collaborators

21 papers

cs.CV2026

AGFT: Alignment-Guided Fine-Tuning for Zero-Shot Adversarial Robustness of Vision-Language Models

Yubo Cui, Xianchao Guan, Zijun Xiong +1

Pre-trained vision-language models (VLMs) exhibit strong zero-shot generalization but remain vulnerable to adversarial perturbations. Existing classification-guided adversarial fin…

cs.CV2026

Rényi Entropy: A New Token Pruning Metric for Vision Transformers

Wei-Yuan Su, Ruijie Zhang, Zheng Zhang

Vision Transformers (ViTs) achieve state-of-the-art performance but suffer from the complexity of self-attention, making inference costly for high-resolution inputs. To ad…

cs.CV20231 cited

DETR Doesn't Need Multi-Scale or Locality Design

Yutong Lin, Yuhui Yuan, Zheng Zhang +3

This paper presents an improved DETR detector that maintains a "plain" nature: using a single-scale feature map and global cross-attention calculations without specific locality co…

cs.CV202212 cited

Exploring Discrete Diffusion Models for Image Captioning

Zixin Zhu, Yixuan Wei, Jianfeng Wang +7

The image captioning task is typically realized by an auto-regressive method that decodes the text tokens one by one. We present a diffusion-based captioning model, dubbed the name…

cs.CV20222 cited

Could Giant Pretrained Image Models Extract Universal Representations?

Yutong Lin, Ze Liu, Zheng Zhang +4

Frozen pretrained models have become a viable alternative to the pretraining-then-finetuning paradigm for transfer learning. However, with frozen models there are relatively few pa…

cs.CV20224 cited

Expediting Large-Scale Vision Transformer for Dense Prediction without Fine-tuning

Weicong Liang, Yuhui Yuan, Henghui Ding +6

Vision transformers have recently achieved competitive results across various vision tasks but still suffer from heavy computation costs when processing a large number of tokens. M…