papers

Publications (12)

cs.LG2026

TabPFN-3: Technical Report

Léo Grinsztajn, Klemens Flöge, Oscar Key +38

Tabular data underpins most high-value prediction problems in science and industry, and TabPFN has driven the foundation model revolution for this modality. Designed with feedback…

stat.ML2021

On Signal-to-Noise Ratio Issues in Variational Inference for Deep Gaussian Processes

Tim G. J. Rudner, Oscar Key, Yarin Gal +1

We show that the gradient estimates used in training Deep Gaussian Processes (DGPs) with importance-weighted variational inference are susceptible to signal-to-noise ratio (SNR) is…

cs.LG2026

TabPFN-2.5: Advancing the State of the Art in Tabular Foundation Models

Léo Grinsztajn, Klemens Flöge, Oscar Key +23

The first tabular foundation model, TabPFN, and its successor TabPFNv2 have impacted tabular AI substantially, with dozens of methods building on it and hundreds of applications ac…

stat.ML2025

Composite Goodness-of-fit Tests with Kernels

Oscar Key, Arthur Gretton, François-Xavier Briol +1

Model misspecification can create significant challenges for the implementation of probabilistic models, and this has led to development of a range of robust methods which directly…

cs.LG2024

Scalable Data Assimilation with Message Passing

Oscar Key, So Takao, Daniel Giles +1

Data assimilation is a core component of numerical weather prediction systems. The large quantity of data processed during assimilation requires the computation to be distributed a…

stat.ME2023

Optimally-Weighted Estimators of the Maximum Mean Discrepancy for Likelihood-Free Inference

Ayush Bharti, Masha Naslidnyk, Oscar Key +2

Likelihood-free inference methods typically make use of a distance between simulated and real data. A common example is the maximum mean discrepancy (MMD), which has previously bee…

cs.LG2022

Interlocking Backpropagation: Improving depthwise model-parallelism

Aidan N. Gomez, Oscar Key, Kuba Perlin +4

The number of parameters in state of the art neural networks has drastically increased in recent years. This surge of interest in large scale neural networks has motivated the deve…

cs.LG2024

Approximate Top- for Increased Parallelism

Oscar Key, Luka Ribar, Alberto Cattaneo +2

We present an evaluation of bucketed approximate top- algorithms. Computing top- exactly suffers from limited parallelism, because the largest values must be aggregated a…

stat.ML2025

Towards Healing the Blindness of Score Matching

Mingtian Zhang, Oscar Key, Peter Hayes +3

Score-based divergences have been widely used in machine learning and statistics applications. Despite their empirical success, a blindness problem has been observed when using the…

cs.LG2022

On Feature Collapse and Deep Kernel Learning for Single Forward Pass Uncertainty

Joost van Amersfoort, Lewis Smith, Andrew Jesson +2

Inducing point Gaussian process approximations are often considered a gold standard in uncertainty estimation since they retain many of the properties of the exact GP and scale to…

cs.LG2021

Generating Interpretable Counterfactual Explanations By Implicit Minimisation of Epistemic and Aleatoric Uncertainties

Lisa Schut, Oscar Key, Rory McGrath +4

Counterfactual explanations (CEs) are a practical tool for demonstrating why machine learning classifiers make particular decisions. For CEs to be useful, it is important that they…

cs.LG2023

No Train No Gain: Revisiting Efficient Training Algorithms For Transformer-based Language Models

Jean Kaddour, Oscar Key, Piotr Nawrot +2

The computation necessary for training Transformer-based language models has skyrocketed in recent years. This trend has motivated research on efficient training algorithms designe…