113 citations · 124 across the 4 of their papers we have counts for
9 papers
The Privacy Power of Correlated Noise in Decentralized Learning
Youssef Allouah, Anastasia Koloskova, Aymane El Firdoussi +2
Decentralized learning is appealing as it enables the scalable usage of large amounts of distributed data and resources (without resorting to any central entity), while promoting p…
DenseFormer: Enhancing Information Flow in Transformers via Depth Weighted Averaging
Matteo Pagliardini, Amirkeivan Mohtashami, Francois Fleuret +1
The transformer architecture by Vaswani et al. (2017) is now ubiquitous across application domains, from natural language processing to speech processing and image understanding. W…
Controllable Topic-Focused Abstractive Summarization
Seyed Ali Bahrainian, Martin Jaggi, Carsten Eickhoff
Controlled abstractive summarization focuses on producing condensed versions of a source article to cover specific aspects by shifting the distribution of generated text towards a…
Irreducible Curriculum for Language Model Pretraining
Simin Fan, Martin Jaggi
Automatic data selection and curriculum design for training large language models is challenging, with only a few existing methods showing improvements over standard training. Furt…
MultiModN- Multimodal, Multi-Task, Interpretable Modular Networks
Vinitra Swamy, Malika Satayeva, Jibril Frej +5
Predicting multiple real-world tasks in a single model often requires a particularly diverse feature space. Multimodal (MM) models aim to extract the synergistic predictive potenti…
Faster Causal Attention Over Large Sequences Through Sparse Flash Attention
Matteo Pagliardini, Daniele Paliotta, Martin Jaggi +1
Transformer-based language models have found many diverse applications requiring them to process sequences of increasing length. For these applications, the causal self-attention -…