3 papers
cs.LG2026
BAITBENCH: Measuring Agent Reward Hacking with Optional Shortcuts Planted in ML Tasks
Pradyumna Shyama Prasad, Meiri Anto, Leon Eshuijs +3
LLM agents are increasingly used to run autonomous ML experiments, iterating on target metrics with little human oversight. Prior work has documented reward hacking in these enviro…
cs.LG2026
Soft Contamination Means Benchmarks Test Shallow Generalization
Ari Spiesberger, Juan J. Vazquez, Nicky Pochinkov +4
If LLM training data is polluted with benchmark test data, then benchmark performance gives biased estimates of out-of-distribution (OOD) generalization. Typical decontamination fi…
cs.LG2024
Questionable practices in machine learning
Gavin Leech, Juan J. Vazquez, Niclas Kupper +2
Evaluating modern ML models is hard. The strong incentive for researchers and companies to report a state-of-the-art result on some metric often leads to questionable research prac…