Publications (12)
TabPFN-3: Technical Report
Léo Grinsztajn, Klemens Flöge, Oscar Key +38
Tabular data underpins most high-value prediction problems in science and industry, and TabPFN has driven the foundation model revolution for this modality. Designed with feedback…
On Signal-to-Noise Ratio Issues in Variational Inference for Deep Gaussian Processes
Tim G. J. Rudner, Oscar Key, Yarin Gal +1
We show that the gradient estimates used in training Deep Gaussian Processes (DGPs) with importance-weighted variational inference are susceptible to signal-to-noise ratio (SNR) is…
TabPFN-2.5: Advancing the State of the Art in Tabular Foundation Models
Léo Grinsztajn, Klemens Flöge, Oscar Key +23
The first tabular foundation model, TabPFN, and its successor TabPFNv2 have impacted tabular AI substantially, with dozens of methods building on it and hundreds of applications ac…
Composite Goodness-of-fit Tests with Kernels
Oscar Key, Arthur Gretton, François-Xavier Briol +1
Model misspecification can create significant challenges for the implementation of probabilistic models, and this has led to development of a range of robust methods which directly…
Scalable Data Assimilation with Message Passing
Oscar Key, So Takao, Daniel Giles +1
Data assimilation is a core component of numerical weather prediction systems. The large quantity of data processed during assimilation requires the computation to be distributed a…
Optimally-Weighted Estimators of the Maximum Mean Discrepancy for Likelihood-Free Inference
Ayush Bharti, Masha Naslidnyk, Oscar Key +2
Likelihood-free inference methods typically make use of a distance between simulated and real data. A common example is the maximum mean discrepancy (MMD), which has previously bee…
Interlocking Backpropagation: Improving depthwise model-parallelism
Aidan N. Gomez, Oscar Key, Kuba Perlin +4
The number of parameters in state of the art neural networks has drastically increased in recent years. This surge of interest in large scale neural networks has motivated the deve…
Approximate Top- for Increased Parallelism
Oscar Key, Luka Ribar, Alberto Cattaneo +2
We present an evaluation of bucketed approximate top- algorithms. Computing top- exactly suffers from limited parallelism, because the largest values must be aggregated a…
Towards Healing the Blindness of Score Matching
Mingtian Zhang, Oscar Key, Peter Hayes +3
Score-based divergences have been widely used in machine learning and statistics applications. Despite their empirical success, a blindness problem has been observed when using the…
On Feature Collapse and Deep Kernel Learning for Single Forward Pass Uncertainty
Joost van Amersfoort, Lewis Smith, Andrew Jesson +2
Inducing point Gaussian process approximations are often considered a gold standard in uncertainty estimation since they retain many of the properties of the exact GP and scale to…
Generating Interpretable Counterfactual Explanations By Implicit Minimisation of Epistemic and Aleatoric Uncertainties
Lisa Schut, Oscar Key, Rory McGrath +4
Counterfactual explanations (CEs) are a practical tool for demonstrating why machine learning classifiers make particular decisions. For CEs to be useful, it is important that they…
No Train No Gain: Revisiting Efficient Training Algorithms For Transformer-based Language Models
Jean Kaddour, Oscar Key, Piotr Nawrot +2
The computation necessary for training Transformer-based language models has skyrocketed in recent years. This trend has motivated research on efficient training algorithms designe…