56 citations · 82 across the 3 of their papers we have counts for
7 papers
ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate
Chi-Min Chan, Weize Chen, Yusheng Su +5
Text evaluation has historically posed significant challenges, often demanding substantial labor and time cost. With the emergence of large language models (LLMs), researchers have…
ChatDB: Augmenting LLMs with Databases as Their Symbolic Memory
Chenxu Hu, Jie Fu, Chenzhuang Du +3
Large language models (LLMs) with memory are computationally universal. However, mainstream LLMs are not taking full advantage of memory, and the designs are heavily influenced by…
Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs
Miao Xiong, Zhiyuan Hu, Xinyang Lu +4
Empowering large language models to accurately express confidence in their answers is essential for trustworthy decision-making. Previous confidence elicitation methods, which prim…
Huatuo-26M, a Large-scale Chinese Medical QA Dataset
Jianquan Li, Xidong Wang, Xiangbo Wu +6
In this paper, we release a largest ever medical Question Answering (QA) dataset with 26 million QA pairs. We benchmark many existing approaches in our dataset in terms of both ret…
Med-UniC: Unifying Cross-Lingual Medical Vision-Language Pre-Training by Diminishing Bias
Zhongwei Wan, Che Liu, Mi Zhang +6
The scarcity of data presents a critical obstacle to the efficacy of medical visionlanguage pre-training (VLP). A potential solution lies in the combination of datasets from variou…
Assessing Hidden Risks of LLMs: An Empirical Study on Robustness, Consistency, and Credibility
Wentao Ye, Mingfeng Ou, Tianyi Li +8
The recent popularity of large language models (LLMs) has brought a significant impact to boundless fields, particularly through their open-ended ecosystem such as the APIs, open-s…