2 papers
stat.ML2026
Towards Reliable LLM Evaluation: Correcting the Winner's Curse in Adaptive Benchmarking
Yang Xu, Jiefu Zhang, Haixiang Sun +3
Adaptive prompt and program search makes LLM evaluation selection-sensitive. Once benchmark items are reused inside tuning, the observed winner's score need not estimate the fresh-…
cs.DC2026
SkyNomad: On Using Multi-Region Spot Instances to Minimize AI Batch Job Cost
Zhifei Li, Tian Xia, Ziming Mao +9
AI batch jobs such as model training, inference pipelines, and data analytics require substantial GPU resources and often need to finish before a deadline. Spot instances offer 3-1…