8 citations · 12 across the 3 of their papers we have counts for
3 papers
cs.CL2024★ 2 cited
MedBench: A Comprehensive, Standardized, and Reliable Benchmarking System for Evaluating Chinese Medical Large Language Models
Mianxin Liu, Jinru Ding, Jie Xu +16
Ensuring the general efficacy and goodness for human beings from medical large language models (LLM) before real-world deployment is crucial. However, a widely accepted and accessi…
cs.CV2024★ 8 cited
OpenMEDLab: An Open-source Platform for Multi-modality Foundation Models in Medicine
Xiaosong Wang, Xiaofan Zhang, Guotai Wang +17
The emerging trend of advancing generalist artificial intelligence, such as GPTv4 and Gemini, has reshaped the landscape of research (academia and industry) in machine learning and…
cs.CL2023★ 2 cited
MedGPTEval: A Dataset and Benchmark to Evaluate Responses of Large Language Models in Medicine
Jie Xu, Lu Lu, Sen Yang +10
METHODS: First, a set of evaluation criteria is designed based on a comprehensive literature review. Second, existing candidate criteria are optimized for using a Delphi method by…