8 papers
Rotary Position Encodings for Graphs
Isaac Reid, Arijit Sehanobish, Cederik Höfs +7
We study the extent to which rotary position encodings (RoPE), a recent transformer position encoding algorithm broadly adopted in large language models (LLMs) and vision transform…
Incremental Transformer Neural Processes
Philip Mortimer, Cristiana Diaconu, Tommy Rochussen +2
Neural Processes (NPs), and specifically Transformer Neural Processes (TNPs), have demonstrated remarkable performance across tasks ranging from spatiotemporal forecasting to tabul…
Better Hessians Matter: Studying the Impact of Curvature Approximations in Influence Functions
Steve Hong, Runa Eschenhagen, Bruno Mlodozeniec +1
Influence functions offer a principled way to trace model predictions back to training data, but their use in deep learning is hampered by the need to invert a large, ill-condition…
Probabilistic Modelling is Sufficient for Causal Inference
Bruno Mlodozeniec, David Krueger, Richard E. Turner
Causal inference is a key research area in machine learning, yet confusion reigns over the tools needed to tackle it. There are prevalent claims in the machine learning literature…
Distributional Training Data Attribution: What do Influence Functions Sample?
Bruno Mlodozeniec, Isaac Reid, Sam Power +4
Randomness is an unavoidable part of training deep learning models, yet something that traditional training data attribution algorithms fail to rigorously account for. They ignore…
Influence Functions for Scalable Data Attribution in Diffusion Models
Bruno Mlodozeniec, Runa Eschenhagen, Juhan Bae +3
Diffusion models have led to significant advancements in generative modelling. Yet their widespread adoption poses challenges regarding data attribution and interpretability. In th…