activity
20192024
most citedLearning Layer-wise Equivariances Automatically using Gradients

2 citations · 4 across the 8 of their papers we have counts for

collaborators
Showing cs.LGShow all

11 papers · 1 filter

cs.LG2024

Uncertainty-Penalized Direct Preference Optimization

Sam Houliston, Alizée Pace, Alexander Immer +1

Aligning Large Language Models (LLMs) to human preferences in content, style, and presentation is challenging, in part because preferences are varied, context-dependent, and someti…

cs.LG2024

Influence Functions for Scalable Data Attribution in Diffusion Models

Bruno Mlodozeniec, Runa Eschenhagen, Juhan Bae +3

Diffusion models have led to significant advancements in generative modelling. Yet their widespread adoption poses challenges regarding data attribution and interpretability. In th…

cs.LG2024

Shaving Weights with Occam's Razor: Bayesian Sparsification for Neural Networks Using the Marginal Likelihood

Rayen Dhahri, Alexander Immer, Betrand Charpentier +2

Neural network sparsification is a promising avenue to save computational time and memory costs, especially in an age where many successful AI models are becoming too large to naïv…

cs.LG2024

Position: Bayesian Deep Learning is Needed in the Age of Large-Scale AI

Theodore Papamarkou, Maria Skoularidou, Konstantina Palla +22

In the current landscape of deep learning research, there is a predominant emphasis on achieving high predictive accuracy in supervised tasks involving large image and language dat…

cs.LG2023

Uncertainty in Graph Contrastive Learning with Bayesian Neural Networks

Alexander Möllers, Alexander Immer, Elvin Isufi +1

Graph contrastive learning has shown great promise when labeled data is scarce, but large unlabeled datasets are available. However, it often does not take uncertainty estimation i…

cs.LG2023

Kronecker-Factored Approximate Curvature for Modern Neural Network Architectures

Runa Eschenhagen, Alexander Immer, Richard E. Turner +2

The core components of many modern neural network architectures, such as transformers, convolutional, or graph neural networks, can be expressed as linear layers with $\textit{weig…