2 papers
cs.LG2026
Sparser, Faster, Lighter Transformer Language Models
Edoardo Cetin, Stefano Peluchetti, Emilio Castillo +3
Scaling autoregressive large language models (LLMs) has driven unprecedented progress but comes with vast computational costs. In this work, we tackle these costs by leveraging uns…
cs.CL2025
Transformer Layers as Painters
Qi Sun, Marc Pickett, Aakash Kumar Nain +1
Despite their nearly universal adoption for large language models, the internal workings of transformers are not well understood. We aim to better understand the impact of removing…