Showing cs.LGShow all
2 papers · 1 filter
cs.LG2024★ 4 cited
Mixture-of-Depths: Dynamically allocating compute in transformer-based language models
David Raposo, Sam Ritter, Blake Richards +3
Transformer-based language models spread FLOPs uniformly across input sequences. In this work we demonstrate that transformers can instead learn to dynamically allocate FLOPs (or c…
cs.LG2023★ 22 cited
A Unified, Scalable Framework for Neural Population Decoding
Mehdi Azabou, Vinam Arora, Venkataramana Ganesh +7
Our ability to use deep learning approaches to decipher neural activity would likely benefit from greater scale, in terms of both model size and datasets. However, the integration…