4 citations · 6 across the 5 of their papers we have counts for
5 papers
AbbIE: Autoregressive Block-Based Iterative Encoder for Efficient Sequence Modeling
Preslav Aleksandrov, Meghdad Kurmanji, Fernando Garcia Redondo +7
We introduce the Autoregressive Block-Based Iterative Encoder (AbbIE), a novel recursive generalization of the encoder-only Transformer architecture, which achieves better perplexi…
DES-LOC: Desynced Low Communication Adaptive Optimizers for Training Foundation Models
Alex Iacob, Lorenzo Sani, Mher Safaryan +8
Scaling foundation model training with Distributed Data Parallel (DDP) methods is bandwidth-limited. Existing infrequent communication methods like Local SGD were designed to synch…
Compliance Cards: Automated EU AI Act Compliance Analyses amidst a Complex AI Supply Chain
Bill Marino, Yaqub Chaudhary, Yulu Pi +6
As the AI supply chain grows more complex, AI systems and models are increasingly likely to incorporate multiple internally- or externally-sourced components such as datasets and (…
Worldwide Federated Training of Language Models
Alex Iacob, Lorenzo Sani, Bill Marino +3
The reliance of language model training on massive amounts of computation and vast datasets scraped from potentially low-quality, copyrighted, or sensitive data has come into quest…
The Future of Large Language Model Pre-training is Federated
Lorenzo Sani, Alex Iacob, Zeyu Cao +8
Generative pre-trained large language models (LLMs) have demonstrated impressive performance over a wide range of tasks, thanks to the unprecedented amount of data they have been t…