2 papers
stat.ME2026
Scalable Text-Embedding-informed Cognitive Diagnosis of Large Language Models
Jia Liu, Zhiyu Xu, Yuqi Gu
Large language models (LLMs) have achieved remarkable performance on diverse benchmarks, yet existing evaluation practices largely rely on coarse summary metrics that obscure under…
stat.ME2025
Latency-Response Theory Model: Evaluating Large Language Models via Response Accuracy and Chain-of-Thought Length
Zhiyu Xu, Jia Liu, Yixin Wang +1
The proliferation of Large Language Models (LLMs) necessitates valid evaluation methods to provide guidance for both downstream applications and actionable future improvements. The…