2 citations · 3 across the 4 of their papers we have counts for
6 papers
When Does Sparsity Mitigate the Curse of Depth in LLMs
Dilxat Muhtar, Xinyuan Song, Sebastian Pokutta +4
Recent work has demonstrated the curse of depth in large language models (LLMs), where later layers contribute less to learning and representation than earlier layers. Such under-u…
A Language Model's Guide Through Latent Space
Dimitri von Rütte, Sotiris Anagnostidis, Gregor Bachmann +1
Concept guidance has emerged as a cheap and simple way to control the behavior of language models by probing their hidden representations for concept vectors and using them to pert…
How Good is a Single Basin?
Kai Lion, Lorenzo Noci, Thomas Hofmann +1
The multi-modal nature of neural loss landscapes is often considered to be the main driver behind the empirical success of deep ensembles. In this work, we probe this belief by con…
Disentangling Linear Mode-Connectivity
Gul Sena Altintas, Gregor Bachmann, Lorenzo Noci +1
Linear mode-connectivity (LMC) (or lack thereof) is one of the intriguing characteristics of neural network loss landscapes. While empirically well established, it unfortunately st…
Harnessing Synthetic Datasets: The Role of Shape Bias in Deep Neural Network Generalization
Elior Benarous, Sotiris Anagnostidis, Luca Biggio +1
Recent advancements in deep learning have been primarily driven by the use of large models trained on increasingly vast datasets. While neural scaling laws have emerged to predict…
Navigating Scaling Laws: Compute Optimality in Adaptive Model Training
Sotiris Anagnostidis, Gregor Bachmann, Imanol Schlag +1
In recent years, the state-of-the-art in deep learning has been dominated by very large models that have been pre-trained on vast amounts of data. The paradigm is very simple: inve…