8 papers · 1 filter
Efficient Adversarial Training via Criticality-Aware Fine-Tuning
Wenyun Li, Zheng Zhang, Dongmei Jiang +2
Vision Transformer (ViT) models have achieved remarkable performance across various vision tasks, with scalability being a key advantage when applied to large datasets. This scalab…
Toward Visual Grounding: A Survey
Linhui Xiao, Xiaoshan Yang, Xiangyuan Lan +2
Visual Grounding, also known as Referring Expression Comprehension and Phrase Grounding, aims to ground the specific region(s) within the image(s) based on the given expression tex…
SelaVPR++: Towards Seamless Adaptation of Foundation Models for Efficient Place Recognition
Feng Lu, Tong Jin, Xiangyuan Lan +4
Recent studies show that the visual place recognition (VPR) method using pre-trained visual foundation models can achieve promising performance. In our previous work, we propose a…
DS-Det: Single-Query Paradigm and Attention Disentangled Learning for Flexible Object Detection
Guiping Cao, Xiangyuan Lan, Wenjian Huang +3
Popular transformer detectors have achieved promising performance through query-based learning using attention mechanisms. However, the roles of existing decoder query types (e.g.,…
Cross-DINO: Cross the Deep MLP and Transformer for Small Object Detection
Guiping Cao, Wenjian Huang, Xiangyuan Lan +3
Small Object Detection (SOD) poses significant challenges due to limited information and the model's low class prediction score. While Transformer-based detectors have shown promis…
CATCH: Complementary Adaptive Token-level Contrastive Decoding to Mitigate Hallucinations in LVLMs
Zhehan Kan, Ce Zhang, Zihan Liao +7
Large Vision-Language Model (LVLM) systems have demonstrated impressive vision-language reasoning capabilities but suffer from pervasive and severe hallucination issues, posing sig…