4 papers · 1 filter
CASTLE: A Comprehensive Benchmark for Evaluating Student-Tailored Personalized Safety in Large Language Models
Rui Jia, Ruiyi Lan, Fengrui Liu +7
Large language models (LLMs) have advanced the development of personalized learning in education. However, their inherent generation mechanisms often produce homogeneous responses…
Reversible Diffusion Decoding for Diffusion Language Models
Xinyun Wang, Min Zhang, Sen Cui +4
Diffusion language models enable parallel token generation through block-wise decoding, but their irreversible commitments can lead to stagnation, where the reverse diffusion proce…
OmniEduBench: A Comprehensive Chinese Benchmark for Evaluating Large Language Models in Education
Min Zhang, Hao Chen, Wenqi Zhang +6
With the rapid development of large language models (LLMs), various LLM-based works have been widely applied in educational fields. However, most existing LLMs and their benchmarks…
EduDial: Constructing a Large-scale Multi-turn Teacher-Student Dialogue Corpus
Shouang Wei, Min Zhang, Xin Lin +3
Recently, several multi-turn dialogue benchmarks have been proposed to evaluate the conversational abilities of large language models (LLMs). As LLMs are increasingly recognized as…