2 papers
cs.CV2026
The Model Knows Which Tokens Matter: Automatic Token Selection via Noise Gating
Landi He, Xiaoyu Yang, Lijian Xu
Visual tokens dominate inference cost in vision-language models (VLMs), yet many carry redundant information. Existing pruning methods alleviate this but typically rely on attentio…
cs.CV2025
One Leaf Reveals the Season: Occlusion-Based Contrastive Learning with Semantic-Aware Views for Efficient Visual Representation
Xiaoyu Yang, Lijian Xu, Hongsheng Li +1
This paper proposes a scalable and straightforward pre-training paradigm for efficient visual conceptual representation called occluded image contrastive learning (OCL). Our OCL ap…