39 citations · 83 across the 19 of their papers we have counts for
6 papers · 2 filters
Patch is Enough: Naturalistic Adversarial Patch against Vision-Language Pre-training Models
Dehong Kong, Siyuan Liang, Xiaopeng Zhu +2
Visual language pre-training (VLP) models have demonstrated significant success across various domains, yet they remain vulnerable to adversarial attacks. Addressing these adversar…
Uncertainty-aware sign language video retrieval with probability distribution modeling
Xuan Wu, Hongxiang Li, Yuanjiang Luo +4
Sign language video retrieval plays a key role in facilitating information access for the deaf community. Despite significant advances in video-text retrieval, the complexity and i…
Effectiveness Assessment of Recent Large Vision-Language Models
Yao Jiang, Xinyu Yan, Ge-Peng Ji +5
The advent of large vision-language models (LVLMs) represents a remarkable advance in the quest for artificial general intelligence. However, the model's effectiveness in both spec…
Explicit Motion Handling and Interactive Prompting for Video Camouflaged Object Detection
Xin Zhang, Tao Xiao, Gepeng Ji +3
Camouflage poses challenges in distinguishing a static target, whereas any movement of the target can break this disguise. Existing video camouflaged object detection (VCOD) approa…
Dynamic Patch-aware Enrichment Transformer for Occluded Person Re-Identification
Xin Zhang, Keren Fu, Qijun Zhao
Person re-identification (re-ID) continues to pose a significant challenge, particularly in scenarios involving occlusions. Prior approaches aimed at tackling occlusions have predo…
Promoting Segment Anything Model towards Highly Accurate Dichotomous Image Segmentation
Xianjie Liu, Keren Fu, Yao Jiang +1
The Segment Anything Model (SAM) represents a significant breakthrough into foundation models for computer vision, providing a large-scale image segmentation model. However, despit…