3 papers
cs.LG2026
CapBencher: Give Your LLM Benchmark a Built-in Alarm for Test-Set Overfitting
Takashi Ishida, Thanawat Lodkaew, Ikko Yamane
Publishing a large language model (LLM) benchmark (especially its ground-truth answers) on the Internet risks contaminating future LLMs and enabling evaluation gaming: it may be un…
math.ST2025
Convergence rate for Nearest Neighbour matching: geometry of the domain and higher-order regularity
Simon Viel, Lionel Truquet, Ikko Yamane
Estimating some mathematical expectations from partially observed data and in particular missing outcomes is a central problem encountered in numerous fields such as transfer learn…
stat.ML2024
Nearest Neighbor Sampling for Covariate Shift Adaptation
François Portier, Lionel Truquet, Ikko Yamane
Many existing covariate shift adaptation methods estimate sample weights given to loss values to mitigate the gap between the source and the target distribution. However, estimatin…