3 papers
cs.LG2026
Sparser, Faster, Lighter Transformer Language Models
Edoardo Cetin, Stefano Peluchetti, Emilio Castillo +3
Scaling autoregressive large language models (LLMs) has driven unprecedented progress but comes with vast computational costs. In this work, we tackle these costs by leveraging uns…
cs.LG2024
The Ungrounded Alignment Problem
Marc Pickett, Aakash Kumar Nain, Joseph Modayil +1
Modern machine learning systems have demonstrated substantial abilities with methods that either embrace or ignore human-provided knowledge, but combining benefits of both styles r…
cs.CL2024
Transformer Layers as Painters
Qi Sun, Marc Pickett, Aakash Kumar Nain +1
Despite their nearly universal adoption for large language models, the internal workings of transformers are not well understood. We aim to better understand the impact of removing…