Showing cs.AIShow all
2 papers · 1 filter
cs.AI2025
am-ELO: A Stable Framework for Arena-based LLM Evaluation
Zirui Liu, Jiatong Li, Yan Zhuang +5
Arena-based evaluation is a fundamental yet significant evaluation paradigm for modern AI models, especially large language models (LLMs). Existing framework based on ELO rating sy…
cs.AI2024
A Survey of Models for Cognitive Diagnosis: New Developments and Future Directions
Fei Wang, Weibo Gao, Qi Liu +8
Cognitive diagnosis has been developed for decades as an effective measurement tool to evaluate human cognitive status such as ability level and knowledge mastery. It has been appl…