8 citations · 11 across the 2 of their papers we have counts for
2 papers
cs.AI2025★ 8 cited
AAKT: Enhancing Knowledge Tracing with Alternate Autoregressive Modeling
Hao Zhou, Wenge Rong, Jianfei Zhang +3
Knowledge Tracing (KT) aims to predict students' future performances based on their former exercises and additional information in educational settings. KT has received significant…
cs.CL2022★ 3 cited
Mixture of Attention Heads: Selecting Attention Heads Per Token
Xiaofeng Zhang, Yikang Shen, Zeyu Huang +3
Mixture-of-Experts (MoE) networks have been proposed as an efficient way to scale up model capacity and implement conditional computing. However, the study of MoE components mostly…