6 papers
Inflated Excellence or True Performance? Rethinking Medical Diagnostic Benchmarks with Dynamic Evaluation
Xiangxu Zhang, Lei Li, Yanyun Zhou +3
Medical diagnostics is a high-stakes and complex domain that is critical to patient care. However, current evaluations of large language models (LLMs) remain limited in capturing k…
R2MED: A Benchmark for Reasoning-Driven Medical Retrieval
Xiangxu Zhang, Lei Li, Xiao Zhou +1
Current medical retrieval benchmarks primarily emphasize lexical or shallow semantic similarity, overlooking the reasoning-intensive demands that are central to clinical decision-m…
Beyond Single Models: Mitigating Multimodal Hallucinations via Adaptive Token Ensemble Decoding
Jinlin Li, Yuran Wang, Yifei Yuan +5
Large Vision-Language Models (LVLMs) have recently achieved impressive results in multimodal tasks such as image captioning and visual question answering. However, they remain pron…
From Retrieval to Generation: Unifying External and Parametric Knowledge for Medical Question Answering
Lei Li, Xiao Zhou, Yingying Zhang +1
Medical question answering (QA) requires extensive access to domain-specific knowledge. A promising direction is to enhance large language models (LLMs) with external knowledge ret…
AutoMIR: Effective Zero-Shot Medical Information Retrieval without Relevance Labels
Lei Li, Xiangxu Zhang, Xiao Zhou +1
Medical information retrieval (MIR) is essential for retrieving relevant medical knowledge from diverse sources, including electronic health records, scientific literature, and med…
Leave No One Behind: Enhancing Diversity While Maintaining Accuracy in Social Recommendation
Lei Li, Xiao Zhou
Social recommendation, which incorporates social connections into recommender systems, has proven effective in improving recommendation accuracy. However, beyond accuracy, diversit…