4 papers
OmniEduBench: A Comprehensive Chinese Benchmark for Evaluating Large Language Models in Education
Min Zhang, Hao Chen, Wenqi Zhang +6
With the rapid development of large language models (LLMs), various LLM-based works have been widely applied in educational fields. However, most existing LLMs and their benchmarks…
DiaCDM: Cognitive Diagnosis in Teacher-Student Dialogues using the Initiation-Response-Evaluation Framework
Rui Jia, Yuang Wei, Ruijia Li +5
While cognitive diagnosis (CD) effectively assesses students' knowledge mastery from structured test data, applying it to real-world teacher-student dialogues presents two fundamen…
ELMES: An Automated Framework for Evaluating Large Language Models in Educational Scenarios
Shou'ang Wei, Xinyun Wang, Shuzhen Bi +9
The emergence of Large Language Models (LLMs) presents transformative opportunities for education, generating numerous novel application scenarios. However, significant challenges…
CALM: Consensus-Aware Localized Merging for Multi-Task Learning
Kunda Yan, Min Zhang, Sen Cui +4
Model merging aims to integrate the strengths of multiple fine-tuned models into a unified model while preserving task-specific capabilities. Existing methods, represented by task…