2 papers
cs.CL2024
PertEval: Unveiling Real Knowledge Capacity of LLMs with Knowledge-Invariant Perturbations
Jiatong Li, Renjun Hu, Kunzhe Huang +5
Expert-designed close-ended benchmarks are indispensable in assessing the knowledge capacity of large language models (LLMs). Despite their widespread use, concerns have mounted re…
cs.LG2024
Survey of Computerized Adaptive Testing: A Machine Learning Perspective
Yan Zhuang, Qi Liu, Haoyang Bi +12
Computerized Adaptive Testing (CAT) offers an efficient and personalized method for assessing examinee proficiency by dynamically adjusting test questions based on individual perfo…