Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025
Tracing the Representation Geometry of Language Models from Pretraining to Post-training
Melody Zixuan Li, Kumar Krishna Agrawal, Arna Ghosh +4
Standard training metrics like loss fail to explain the emergence of complex capabilities in large language models. We take a spectral approach to investigate the geometry of learn…
cs.LG2024
Mixture-of-Depths: Dynamically allocating compute in transformer-based language models
David Raposo, Sam Ritter, Blake Richards +3
Transformer-based language models spread FLOPs uniformly across input sequences. In this work we demonstrate that transformers can instead learn to dynamically allocate FLOPs (or c…