2 papers
cs.AI2026
Interactive Benchmarks
Baoqing Yue, Zihan Zhu, Yutong Han +6
Existing reasoning evaluation paradigms suffer from different limitations: fixed benchmarks are increasingly saturated and vulnerable to contamination, while preference-based evalu…
cs.LG2026
Fault-Tolerant Evaluation for Sample-Efficient Model Performance Estimators
Zihan Zhu, Yanqiu Wu, Qiongkai Xu
In the era of Model-as-a-Service, organizations increasingly rely on third-party AI models for rapid deployment. However, the dynamic nature of emerging AI applications, the contin…