4 papers
VideoExpert: Augmented LLM for Temporal-Sensitive Video Understanding
Henghao Zhao, Ge-Peng Ji, Rui Yan +2
The core challenge in video understanding lies in perceiving dynamic content changes over time. However, multimodal large language models struggle with temporal-sensitive video tas…
Diffusion-Enhanced Test-time Adaptation with Text and Image Augmentation
Chun-Mei Feng, Yuanyang He, Jian Zou +6
Existing test-time prompt tuning (TPT) methods focus on single-modality data, primarily enhancing images and using confidence ratings to filter out inaccurate images. However, whil…
Effectiveness Assessment of Recent Large Vision-Language Models
Yao Jiang, Xinyu Yan, Ge-Peng Ji +5
The advent of large vision-language models (LVLMs) represents a remarkable advance in the quest for artificial general intelligence. However, the model's effectiveness in both spec…
Patch is Enough: Naturalistic Adversarial Patch against Vision-Language Pre-training Models
Dehong Kong, Siyuan Liang, Xiaopeng Zhu +2
Visual language pre-training (VLP) models have demonstrated significant success across various domains, yet they remain vulnerable to adversarial attacks. Addressing these adversar…