Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Soft Contamination Means Benchmarks Test Shallow Generalization
Ari Spiesberger, Juan J. Vazquez, Nicky Pochinkov +4
If LLM training data is polluted with benchmark test data, then benchmark performance gives biased estimates of out-of-distribution (OOD) generalization. Typical decontamination fi…
cs.LG2024
Questionable practices in machine learning
Gavin Leech, Juan J. Vazquez, Niclas Kupper +2
Evaluating modern ML models is hard. The strong incentive for researchers and companies to report a state-of-the-art result on some metric often leads to questionable research prac…