3 papers
cs.CV2025
Efficient Token Compression for Vision Transformer with Spatial Information Preserved
Junzhu Mao, Yang Shen, Jinyang Guo +2
Token compression is essential for reducing the computational and memory requirements of transformer models, enabling their deployment in resource-constrained environments. In this…
cs.CV2025
Twofold Debiasing Enhances Fine-Grained Learning with Coarse Labels
Xin-yang Zhao, Jian Jin, Yang-yang Li +1
The Coarse-to-Fine Few-Shot (C2FS) task is designed to train models using only coarse labels, then leverages a limited number of subclass samples to achieve fine-grained recognitio…
cs.CV2024
Explanatory Instructions: Towards Unified Vision Tasks Understanding and Zero-shot Generalization
Yang Shen, Xiu-Shen Wei, Yifan Sun +6
Computer Vision (CV) has yet to fully achieve the zero-shot task generalization observed in Natural Language Processing (NLP), despite following many of the milestones established…