most citedThe Future of Large Language Model Pre-training is Federated

4 citations · 6 across the 5 of their papers we have counts for

collaborators

5 papers

cs.LG2025

AbbIE: Autoregressive Block-Based Iterative Encoder for Efficient Sequence Modeling

Preslav Aleksandrov, Meghdad Kurmanji, Fernando Garcia Redondo +7

We introduce the Autoregressive Block-Based Iterative Encoder (AbbIE), a novel recursive generalization of the encoder-only Transformer architecture, which achieves better perplexi…

cs.LG2025

DES-LOC: Desynced Low Communication Adaptive Optimizers for Training Foundation Models

Alex Iacob, Lorenzo Sani, Mher Safaryan +8

Scaling foundation model training with Distributed Data Parallel (DDP) methods is bandwidth-limited. Existing infrequent communication methods like Local SGD were designed to synch…

cs.AI2024★ 1 cited

Compliance Cards: Automated EU AI Act Compliance Analyses amidst a Complex AI Supply Chain

Bill Marino, Yaqub Chaudhary, Yulu Pi +6

As the AI supply chain grows more complex, AI systems and models are increasingly likely to incorporate multiple internally- or externally-sourced components such as datasets and (…

cs.LG2024★ 1 cited

Worldwide Federated Training of Language Models

Alex Iacob, Lorenzo Sani, Bill Marino +3

The reliance of language model training on massive amounts of computation and vast datasets scraped from potentially low-quality, copyrighted, or sensitive data has come into quest…

cs.LG2024★ 4 cited

The Future of Large Language Model Pre-training is Federated

Lorenzo Sani, Alex Iacob, Zeyu Cao +8

Generative pre-trained large language models (LLMs) have demonstrated impressive performance over a wide range of tasks, thanks to the unprecedented amount of data they have been t…