4 papers
VPTracker: Global Vision-Language Tracking via Visual Prompt
Jingchao Wang, Kaiwen Zhou, Zhijian Wu +3
Vision-Language Tracking aims to continuously localize objects described by a visual template and a language description. Existing methods, however, are typically limited to local…
The Curse of Depth in Large Language Models
Wenfang Sun, Xinyuan Song, Pengxiang Li +3
In this paper, we introduce the Curse of Depth, a concept that highlights, explains, and addresses the recent observation in modern Large Language Models (LLMs) where nearly half o…
Medusa: Cross-Modal Transferable Adversarial Attacks on Multimodal Medical Retrieval-Augmented Generation
Yingjia Shang, Yi Liu, Huimin Wang +4
With the rapid advancement of retrieval-augmented vision-language models, multimodal medical retrieval-augmented generation (MMed-RAG) systems are increasingly adopted in clinical…
MedREK: Retrieval-Based Editing for Medical LLMs with Key-Aware Prompts
Shujun Xia, Haokun Lin, Yichen Wu +9
LLMs hold great promise for healthcare applications, but the rapid evolution of medical knowledge and errors in training data often cause them to generate outdated or inaccurate in…