From the 1 of 7 linked papers with an AI index.
7 papers
LoMeVQA: A Comprehensive Benchmark for Longitudinal Medical VQA
Zhilin Wu, Zhangkai Ni, Chengmei Yang +4
The paper introduces LoMeVQA, a large benchmark of 206K longitudinal medical visual question answering pairs designed to evaluate temporal reasoning over sequential medical images,…
M-IDoL: Information Decomposition for Modality-Specific and Diverse Representation Learning in Medical Foundation Model
Yihang Liu, Longzhen Yang, Jiaxiong Yang +3
Medical foundation models (MFMs) aim to learn universal representations from multimodal medical images that can generalize effectively to diverse downstream clinical tasks. However…
Visual Neural Decoding via Improved Visual-EEG Semantic Consistency
Hongzhou Chen, Lianghua He, Yihang Liu +3
Visual neural decoding aims to extract and interpret original visual experiences directly from human brain activity. Recent studies have demonstrated the feasibility of decoding vi…
When Token Pruning is Worse than Random: Understanding Visual Token Information in VLLMs
Yahong Wang, Juncheng Wu, Zhangkai Ni +8
Vision Large Language Models (VLLMs) incur high computational costs due to their reliance on hundreds of visual tokens to represent images. While token pruning offers a promising s…
EntropyPrune: Matrix Entropy Guided Visual Token Pruning for Multimodal Large Language Models
Yahong Wang, Juncheng Wu, Zhangkai Ni +6
Multimodal large language models (MLLMs) incur substantial inference cost due to the processing of hundreds of visual tokens per image. Although token pruning has proven effective…
Self-Supervised Anatomical Consistency Learning for Vision-Grounded Medical Report Generation
Longzhen Yang, Zhangkai Ni, Ying Wen +3
Vision-grounded medical report generation aims to produce clinically accurate descriptions of medical images, anchored in explicit visual evidence to improve interpretability and f…