5 papers
Towards World Models in Biomedical Research
Guangyu Wang, Jingkun Yue, Siqi Zhang +19
A central goal of biomedicine is to understand, predict and ultimately control the dynamic mechanisms by which biological systems respond to perturbations, disease progression and…
DOTA: Distributional Test-Time Adaptation of Vision-Language Models
Zongbo Han, Jialong Yang, Guangyu Wang +4
Vision-language foundation models (VLMs), such as CLIP, exhibit remarkable performance across a wide range of tasks. However, deploying these models can be unreliable when signific…
Helping CLIP See Both the Forest and the Trees: A Decomposition and Description Approach
Leyan Xue, Zongbo Han, Guangyu Wang +3
Vision-Language Models (VLMs) like CLIP achieve cross-modal semantic alignment through contrastive learning, exhibiting robust zero-shot generalization. Traditional prompt engineer…
Computational Reasoning of Large Language Models
Haitao Wu, Zongbo Han, Joey Tianyi Zhou +2
With the rapid development and widespread application of Large Language Models (LLMs), multidimensional evaluation has become increasingly critical. However, current evaluations ar…
MedSG-Bench: A Benchmark for Medical Image Sequences Grounding
Jingkun Yue, Siqi Zhang, Zinan Jia +4
Visual grounding is essential for precise perception and reasoning in multimodal large language models (MLLMs), especially in medical imaging domains. While existing medical visual…