3 papers
cs.CV2026
Learning to Perceive "Where": Spatial Pretext Tasks for Robust Self-Supervised Learning
Yang Shen, Yusen Cai, Weronika Hryniewska-Guzik +2
Existing self-supervised learning (SSL) methods primarily learn object-invariant representations but often neglect the spatial structure and relationships among object parts. To ad…
cs.CV2025
Explanatory Instructions: Towards Unified Vision Tasks Understanding and Zero-shot Generalization
Yang Shen, Xiu-Shen Wei, Yifan Sun +6
Computer Vision (CV) has yet to fully achieve the zero-shot task generalization observed in Natural Language Processing (NLP), despite following many of the milestones established…
cs.CV2025
Efficient Token Compression for Vision Transformer with Spatial Information Preserved
Junzhu Mao, Yang Shen, Jinyang Guo +2
Token compression is essential for reducing the computational and memory requirements of transformer models, enabling their deployment in resource-constrained environments. In this…