2 papers
cs.AI2026
Agents' Last Exam
Yiyou Sun, Xinyang Han, Weichen Zhang +306
Recent AI systems have achieved strong results on a wide range of benchmarks, yet these gains have not translated into economically meaningful deployment across many professional d…
cs.LG2026
The Post-GCN Decade Revisited: Curvature-Stratified Evaluation of Relational Learning
Shuo Wang, Xiangyu Wang, Quanxin Wang +9
Current evaluation practices in relational learning rely heavily on flat leaderboards that average performance across heterogeneous datasets, implicitly assuming a uniform underlyi…