137 citations · 179 across the 7 of their papers we have counts for
4 papers · 1 filter
Incorporating Cognitive Load and Knowledge Transfer for Multi-Domain Knowledge Tracing
Haotian Zhang, Shucun Wang, Jinze Wu +6
Knowledge Tracing (KT) aims to assess students' dynamic knowledge states from their learning histories. While most existing KT methods focus on single-domain learning with notable…
CogMath: Assessing LLMs' Authentic Mathematical Ability from a Human Cognitive Perspective
Jiayu Liu, Zhenya Huang, Wei Dai +7
Although large language models (LLMs) show promise in solving complex mathematical tasks, existing evaluation paradigms rely solely on a coarse measure of overall answer accuracy,…
am-ELO: A Stable Framework for Arena-based LLM Evaluation
Zirui Liu, Jiatong Li, Yan Zhuang +5
Arena-based evaluation is a fundamental yet significant evaluation paradigm for modern AI models, especially large language models (LLMs). Existing framework based on ELO rating sy…
Quality meets Diversity: A Model-Agnostic Framework for Computerized Adaptive Testing
Haoyang Bi, Haiping Ma, Zhenya Huang +5
Computerized Adaptive Testing (CAT) is emerging as a promising testing application in many scenarios, such as education, game and recruitment, which targets at diagnosing the knowl…