5 papers · 1 filter
Evaluating the Robustness of Chinchilla Compute-Optimal Scaling
Rylan Schaeffer, Noam Levi, Andreas Kirsch +4
Hoffman et al (2022)'s Chinchilla paper introduced the principle of compute-optimal scaling, laying a foundation for future scaling of language models. In the years since, however,…
When three experiments are better than two: Avoiding intractable correlated aleatoric uncertainty by leveraging a novel bias--variance tradeoff
Paul Scherer, Andreas Kirsch, Jake P. Taylor-King
Real-world experimental scenarios are characterized by the presence of heteroskedastic aleatoric uncertainty, and this uncertainty can be correlated in batched settings. The bias--…
(Implicit) Ensembles of Ensembles: Epistemic Uncertainty Collapse in Large Models
Andreas Kirsch
Epistemic uncertainty is crucial for safety-critical applications and data acquisition tasks. Yet, we find an important phenomenon in deep learning models: an epistemic uncertainty…
CoLoR-Filter: Conditional Loss Reduction Filtering for Targeted Language Model Pre-training
David Brandfonbrener, Hanlin Zhang, Andreas Kirsch +2
Selecting high-quality data for pre-training is crucial in shaping the downstream task performance of language models. A major challenge lies in identifying this optimal subset, a…
All models are wrong, some are useful: Model Selection with Limited Labels
Patrik Okanovic, Andreas Kirsch, Jannes Kasper +3
We introduce MODEL SELECTOR, a framework for label-efficient selection of pretrained classifiers. Given a pool of unlabeled target data, MODEL SELECTOR samples a small subset of hi…