1 citations · 1 across the 1 of their papers we have counts for
2 papers
cs.CL2024★ 1 cited
Memory Layers at Scale
Vincent-Pierre Berges, Barlas Oğuz, Daniel Haziza +3
Memory layers use a trainable key-value lookup mechanism to add extra parameters to a model without increasing FLOPs. Conceptually, sparsely activated memory layers complement comp…
cs.CL2024
Byte Latent Transformer: Patches Scale Better Than Tokens
Artidoro Pagnoni, Ram Pasunuru, Pedro Rodriguez +11
We introduce the Byte Latent Transformer (BLT), a new byte-level LLM architecture that, for the first time, matches tokenization-based LLM performance at scale with significant imp…