2 papers
cs.LG2026
Selection-Aware Stress Testing for Interactive Agents
Yang Xu, Chenang Li, Jiefu Zhang +3
Agent evaluations often use one benchmark to choose a workflow and then search for task types where its advantage weakens, so both conclusions are selected from the same data. We i…
stat.ML2026
Towards Reliable LLM Evaluation: Correcting the Winner's Curse in Adaptive Benchmarking
Yang Xu, Jiefu Zhang, Haixiang Sun +3
Adaptive prompt and program search makes LLM evaluation selection-sensitive. Once benchmark items are reused inside tuning, the observed winner's score need not estimate the fresh-…