3 papers
cs.CL2026
LLM Evaluation on Unseen Questions: Contextual Multidimensional IRT Model
Ergan Shang, Weijing Tang, Yinqiu He
Evaluation of large language models (LLMs) increasingly requires predicting how a model will perform on new questions or tasks before collecting large amounts of new annotations. T…
cs.LG2026
ERASE: EaRly bAckpropagation SchEdule for Faster Training of Modern Recommendation Systems
Ergan Shang, Flavio Sales Truzzi
Lightweight proxy models enable rapid experimentation without repeatedly training frontier-scale systems, but their small kernels often leave modern accelerators underutilized. Con…
cs.AI2026
ASI-Bench: At the Dawn of Artificial Superintelligence
Junwei Zhou, Zhen Sun, Binyu Li +39
Artificial superintelligence (ASI) requires AI to move beyond mastering existing knowledge toward exploring the unknown, creating new knowledge, and turning new ideas into verifiab…