Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
Large Language Models Could Be Rote Learners
Yuyang Xu, Renjun Hu, Haochao Ying +3
Benchmark-based evaluation, e.g., multiple-choice questions (MCQs) and open-ended questions (OEQs), is widely used for evaluating Large Language Models (LLMs), yet their reliabilit…
cs.CL2025
Training an LLM-as-a-Judge Model: Pipeline, Insights, and Practical Lessons
Renjun Hu, Yi Cheng, Libin Meng +4
The rapid advancement of large language models (LLMs) has opened new possibilities for their adoption as evaluative judges. This paper introduces Themis, a fine-tuned LLM judge tha…
cs.CL2024
PertEval: Unveiling Real Knowledge Capacity of LLMs with Knowledge-Invariant Perturbations
Jiatong Li, Renjun Hu, Kunzhe Huang +5
Expert-designed close-ended benchmarks are indispensable in assessing the knowledge capacity of large language models (LLMs). Despite their widespread use, concerns have mounted re…