5 papers
Inflated Excellence or True Performance? Rethinking Medical Diagnostic Benchmarks with Dynamic Evaluation
Xiangxu Zhang, Lei Li, Yanyun Zhou +3
Medical diagnostics is a high-stakes and complex domain that is critical to patient care. However, current evaluations of large language models (LLMs) remain limited in capturing k…
Human Values Matter: Investigating How Misalignment Shapes Collective Behaviors in LLM Agent Communities
Xiangxu Zhang, Jiamin Wang, Qinlin Zhao +6
As LLMs become increasingly integrated into human society, evaluating their orientations on human values from social science has drawn growing attention. Nevertheless, it is still…
HypeMed: Enhancing Medication Recommendations with Hypergraph-Based Patient Relationships
Xiangxu Zhang, Xiao Zhou, Hongteng Xu +1
Medication recommendations aim to generate safe and effective medication sets from health records. However, accurately recommending medications hinges on inferring a patient's late…
MoHoBench: Assessing Honesty of Multimodal Large Language Models via Unanswerable Visual Questions
Yanxu Zhu, Shitong Duan, Xiangxu Zhang +7
Recently Multimodal Large Language Models (MLLMs) have achieved considerable advancements in vision-language tasks, yet produce potentially harmful or untrustworthy content. Despit…
AutoMIR: Effective Zero-Shot Medical Information Retrieval without Relevance Labels
Lei Li, Xiangxu Zhang, Xiao Zhou +1
Medical information retrieval (MIR) is essential for retrieving relevant medical knowledge from diverse sources, including electronic health records, scientific literature, and med…