activity
20232026
most citedA Language Model's Guide Through Latent Space

2 citations · 3 across the 4 of their papers we have counts for

collaborators

6 papers

cs.CL2026

When Does Sparsity Mitigate the Curse of Depth in LLMs

Dilxat Muhtar, Xinyuan Song, Sebastian Pokutta +4

Recent work has demonstrated the curse of depth in large language models (LLMs), where later layers contribute less to learning and representation than earlier layers. Such under-u…

cs.CL20242 cited

A Language Model's Guide Through Latent Space

Dimitri von Rütte, Sotiris Anagnostidis, Gregor Bachmann +1

Concept guidance has emerged as a cheap and simple way to control the behavior of language models by probing their hidden representations for concept vectors and using them to pert…

cs.LG2024

How Good is a Single Basin?

Kai Lion, Lorenzo Noci, Thomas Hofmann +1

The multi-modal nature of neural loss landscapes is often considered to be the main driver behind the empirical success of deep ensembles. In this work, we probe this belief by con…

cs.LG2023

Disentangling Linear Mode-Connectivity

Gul Sena Altintas, Gregor Bachmann, Lorenzo Noci +1

Linear mode-connectivity (LMC) (or lack thereof) is one of the intriguing characteristics of neural network loss landscapes. While empirically well established, it unfortunately st…

cs.CV20231 cited

Harnessing Synthetic Datasets: The Role of Shape Bias in Deep Neural Network Generalization

Elior Benarous, Sotiris Anagnostidis, Luca Biggio +1

Recent advancements in deep learning have been primarily driven by the use of large models trained on increasingly vast datasets. While neural scaling laws have emerged to predict…

cs.LG2023

Navigating Scaling Laws: Compute Optimality in Adaptive Model Training

Sotiris Anagnostidis, Gregor Bachmann, Imanol Schlag +1

In recent years, the state-of-the-art in deep learning has been dominated by very large models that have been pre-trained on vast amounts of data. The paradigm is very simple: inve…