most citedChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate

56 citations · 82 across the 3 of their papers we have counts for

collaborators

7 papers

cs.CL202356 cited

ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate

Chi-Min Chan, Weize Chen, Yusheng Su +5

Text evaluation has historically posed significant challenges, often demanding substantial labor and time cost. With the emergence of large language models (LLMs), researchers have…

cs.AI202313 cited

ChatDB: Augmenting LLMs with Databases as Their Symbolic Memory

Chenxu Hu, Jie Fu, Chenzhuang Du +3

Large language models (LLMs) with memory are computationally universal. However, mainstream LLMs are not taking full advantage of memory, and the designs are heavily influenced by…

cs.CL2023

Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs

Miao Xiong, Zhiyuan Hu, Xinyang Lu +4

Empowering large language models to accurately express confidence in their answers is essential for trustworthy decision-making. Previous confidence elicitation methods, which prim…

cs.CL202313 cited

Huatuo-26M, a Large-scale Chinese Medical QA Dataset

Jianquan Li, Xidong Wang, Xiangbo Wu +6

In this paper, we release a largest ever medical Question Answering (QA) dataset with 26 million QA pairs. We benchmark many existing approaches in our dataset in terms of both ret…

cs.CL2023

Med-UniC: Unifying Cross-Lingual Medical Vision-Language Pre-Training by Diminishing Bias

Zhongwei Wan, Che Liu, Mi Zhang +6

The scarcity of data presents a critical obstacle to the efficacy of medical visionlanguage pre-training (VLP). A potential solution lies in the combination of datasets from variou…

cs.LG2023

Assessing Hidden Risks of LLMs: An Empirical Study on Robustness, Consistency, and Credibility

Wentao Ye, Mingfeng Ou, Tianyi Li +8

The recent popularity of large language models (LLMs) has brought a significant impact to boundless fields, particularly through their open-ended ecosystem such as the APIs, open-s…