4 papers · 1 filter
Large Language Models Could Be Rote Learners
Yuyang Xu, Renjun Hu, Haochao Ying +3
Benchmark-based evaluation, e.g., multiple-choice questions (MCQs) and open-ended questions (OEQs), is widely used for evaluating Large Language Models (LLMs), yet their reliabilit…
LLMs Can Simulate Standardized Patients via Agent Coevolution
Zhuoyun Du, Lujie Zheng, Renjun Hu +7
Training medical personnel using standardized patients (SPs) remains a complex challenge, requiring extensive domain expertise and role-specific practice. Previous research on Larg…
Training an LLM-as-a-Judge Model: Pipeline, Insights, and Practical Lessons
Renjun Hu, Yi Cheng, Libin Meng +4
The rapid advancement of large language models (LLMs) has opened new possibilities for their adoption as evaluative judges. This paper introduces Themis, a fine-tuned LLM judge tha…
PertEval: Unveiling Real Knowledge Capacity of LLMs with Knowledge-Invariant Perturbations
Jiatong Li, Renjun Hu, Kunzhe Huang +5
Expert-designed close-ended benchmarks are indispensable in assessing the knowledge capacity of large language models (LLMs). Despite their widespread use, concerns have mounted re…