activity
20092025
most citedLearning Aerial Image Segmentation from Online Maps

284 citations · 1.4k across the 58 of their papers we have counts for

collaborators
Showing cs.LGShow all

76 papers · 1 filter

cs.LG2024★ 3 cited

QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs

Saleh Ashkboos, Amirkeivan Mohtashami, Maximilian L. Croci +6

We introduce QuaRot, a new Quantization scheme based on Rotations, which is able to quantize LLMs end-to-end, including all weights, activations, and KV cache in 4 bits. QuaRot rot…

cs.LG2023★ 2 cited

DoGE: Domain Reweighting with Generalization Estimation

Simin Fan, Matteo Pagliardini, Martin Jaggi

The coverage and composition of the pretraining data significantly impacts the generalization ability of Large Language Models (LLMs). Despite its importance, recent LLMs still rel…

cs.LG2023★ 3 cited

MultiModN- Multimodal, Multi-Task, Interpretable Modular Networks

Vinitra Swamy, Malika Satayeva, Jibril Frej +5

Predicting multiple real-world tasks in a single model often requires a particularly diverse feature space. Multimodal (MM) models aim to extract the synergistic predictive potenti…

cs.LG2023★ 1 cited

Provably Personalized and Robust Federated Learning

Mariel Werner, Lie He, Michael Jordan +2

Identifying clients with similar objectives and learning a model-per-cluster is an intuitive and interpretable approach to personalization in federated learning. However, doing so…

cs.LG2023★ 5 cited

Faster Causal Attention Over Large Sequences Through Sparse Flash Attention

Matteo Pagliardini, Daniele Paliotta, Martin Jaggi +1

Transformer-based language models have found many diverse applications requiring them to process sequences of increasing length. For these applications, the causal self-attention -…

cs.LG2023

On Convergence of Incremental Gradient for Non-Convex Smooth Functions

Anastasia Koloskova, Nikita Doikov, Sebastian U. Stich +1

In machine learning and neural network optimization, algorithms like incremental gradient, and shuffle SGD are popular due to minimizing the number of cache misses and good practic…