1 citations · 1 across the 2 of their papers we have counts for
4 papers · 1 filter
PL-SCEA: Reconfiguring Pretrained Attention for Few-Shot Industrial Anomaly Detection
Xiaoyu Yang, Qixing Wu, Huixian Zhao +1
Vision Foundation Models (VFMs) provide transferable patch representations for few-shot industrial anomaly detection, but their attention computation is typically inherited from pr…
The Model Knows Which Tokens Matter: Automatic Token Selection via Noise Gating
Landi He, Xiaoyu Yang, Lijian Xu
Visual tokens dominate inference cost in vision-language models (VLMs), yet many carry redundant information. Existing pruning methods alleviate this but typically rely on attentio…
One Leaf Reveals the Season: Occlusion-Based Contrastive Learning with Semantic-Aware Views for Efficient Visual Representation
Xiaoyu Yang, Lijian Xu, Hongsheng Li +1
This paper proposes a scalable and straightforward pre-training paradigm for efficient visual conceptual representation called occluded image contrastive learning (OCL). Our OCL ap…
Enhancing Visual Grounding and Generalization: A Multi-Task Cycle Training Approach for Vision-Language Models
Xiaoyu Yang, Lijian Xu, Hao Sun +2
Visual grounding (VG) occupies a pivotal position in multi-modality vision-language models. In this study, we propose ViLaM, a large multi-modality model, that supports multi-tasks…