4 papers
Soft Contamination Means Benchmarks Test Shallow Generalization
Ari Spiesberger, Juan J. Vazquez, Nicky Pochinkov +4
If LLM training data is polluted with benchmark test data, then benchmark performance gives biased estimates of out-of-distribution (OOD) generalization. Typical decontamination fi…
How much technical talent is there? A systematic estimate of the ML research pool among 3 million consultants
Maximilian Schons, Red Bermejo, Florian Aldehoff-Zeidler +4
We identify a substantial pool of technically competent ML research talent (in the low thousands) in companies which offer consulting in machine learning. We systematically searche…
Questionable practices in machine learning
Gavin Leech, Juan J. Vazquez, Niclas Kupper +2
Evaluating modern ML models is hard. The strong incentive for researchers and companies to report a state-of-the-art result on some metric often leads to questionable research prac…
Steering Language Models With Activation Engineering
Alexander Matt Turner, Lisa Thiergart, Gavin Leech +4
Prompt engineering and finetuning aim to maximize language model performance on a given metric (like toxicity reduction). However, these methods do not fully elicit a model's capab…