2 papers
cs.CL2024
CMoralEval: A Moral Evaluation Benchmark for Chinese Large Language Models
Linhao Yu, Yongqi Leng, Yufei Huang +9
What a large language model (LLM) would respond in ethically relevant context? In this paper, we curate a large benchmark CMoralEval for morality evaluation of Chinese LLMs. The da…
cs.AI2024
Establishing Rigorous and Cost-effective Clinical Trials for Artificial Intelligence Models
Wanling Gao, Yunyou Huang, Dandan Cui +22
A profound gap persists between artificial intelligence (AI) and clinical practice in medicine, primarily due to the lack of rigorous and cost-effective evaluation methodologies. S…