4 papers
No Token Left Behind: Efficient Vision Transformer via Dynamic Token Idling
Xuwei Xu, Changlin Li, Yudong Chen +3
Vision Transformers (ViTs) have demonstrated outstanding performance in computer vision tasks, yet their high computational complexity prevents their deployment in computing resour…
Dynamic Token Pruning in Plain Vision Transformers for Semantic Segmentation
Quan Tang, Bowen Zhang, Jiajun Liu +2
Vision transformers have achieved leading performance on various visual tasks yet still suffer from high computational complexity. The situation deteriorates in dense prediction ta…
STAR-GNN: Spatial-Temporal Video Representation for Content-based Retrieval
Guoping Zhao, Bingqing Zhang, Mingyu Zhang +3
We propose a video feature representation learning framework called STAR-GNN, which applies a pluggable graph neural network component on a multi-scale lattice feature graph. The e…
InvisibiliTee: Angle-agnostic Cloaking from Person-Tracking Systems with a Tee
Yaxian Li, Bingqing Zhang, Guoping Zhao +4
After a survey for person-tracking system-induced privacy concerns, we propose a black-box adversarial attack method on state-of-the-art human detection models called InvisibiliTee…