2 papers
cs.CV2025
VACoT: Rethinking Visual Data Augmentation with VLMs
Zhengzhuo Xu, Chong Sun, SiNan Du +3
While visual data augmentation remains a cornerstone for training robust vision models, it has received limited attention in visual language models (VLMs), which predominantly rely…
cs.PF2025
Towards Efficient Multi-Scale Deformable Attention on NPU
Chenghuan Huang, Zhigeng Xu, Chong Sun +2
Multi-scale deformable attention (MSDA) is a flexible and powerful feature extraction mechanism for visual tasks, but its random-access grid sampling strategy poses significant opt…