Publications (60)
VarDiU: A Variational Diffusive Upper Bound for One-Step Diffusion Distillation
Leyang Wang, Mingtian Zhang, Zijing Ou +1
Recently, diffusion distillation methods have compressed thousand-step teacher diffusion models into one-step student generators while preserving sample quality. Most existing appr…
Generating Sentences Using a Dynamic Canvas
Harshil Shah, Bowen Zheng, David Barber
We introduce the Attentive Unsupervised Text (W)riter (AUTR), which is a word level generative model for natural language. It uses a recurrent neural network with a dynamic attenti…
Generative Neural Machine Translation
Harshil Shah, David Barber
We introduce Generative Neural Machine Translation (GNMT), a latent variable architecture which is designed to model the semantics of the source and target sentences. We modify an…
Tracking by Animation: Unsupervised Learning of Multi-Object Attentive Trackers
Zhen He, Jian Li, Daxue Liu +2
Online Multi-Object Tracking (MOT) from videos is a challenging computer vision task which has been extensively studied for decades. Most of the existing MOT algorithms are based o…
Incremental Sequence Classification with Temporal Consistency
Lucas Maystre, Gabriel Barello, Tudor Berariu +5
We address the problem of incremental sequence classification, where predictions are updated as new elements in the sequence are revealed. Drawing on temporal-difference learning f…
Learning disentangled representations with the Wasserstein Autoencoder
Benoit Gaujac, Ilya Feige, David Barber
Disentangled representation learning has undoubtedly benefited from objective function surgery. However, a delicate balancing act of tuning is still required in order to trade off…
Optimal Spike-Timing Dependent Plasticity for Precise Action Potential Firing
Jean-Pascal Pfister, Taro Toyoizumi, David Barber +1
In timing-based neural codes, neurons have to emit action potentials at precise moments in time. We use a supervised learning paradigm to derive a synaptic update rule that optimiz…
Training Neural Samplers with Reverse Diffusive KL Divergence
Jiajun He, Wenlin Chen, Mingtian Zhang +2
Training generative models to sample from unnormalized density functions is an important and challenging task in machine learning. Traditional training methods often rely on the re…
Improving latent variable descriptiveness with AutoGen
Alex Mansbridge, Roberto Fierimonte, Ilya Feige +1
Powerful generative models, particularly in Natural Language Modelling, are commonly trained by maximizing a variational lower bound on the data log likelihood. These models often…
An Approximate Newton Method for Markov Decision Processes
Thomas Furmston, David Barber
Gradient-based algorithms are one of the methods of choice for the optimisation of Markov Decision Processes. In this article we will present a novel approximate Newton algorithm f…
Efficient Inference in Markov Control Problems
Thomas Furmston, David Barber
Markov control algorithms that perform smooth, non-greedy updates of the policy have been shown to be very general and versatile, with policy gradient and Expectation Maximisation…
Nesterov's Accelerated Gradient and Momentum as approximations to Regularised Update Descent
Aleksandar Botev, Guy Lever, David Barber
We present a unifying framework for adapting the update direction in gradient-based iterative optimization methods. As natural special cases we re-derive classical momentum and Nes…
Improving Probabilistic Diffusion Models With Optimal Diagonal Covariance Matching
Zijing Ou, Mingtian Zhang, Andi Zhang +3
The probabilistic diffusion model has become highly effective across various domains. Typically, sampling from a diffusion model involves using a denoising distribution characteriz…
From Characters to Tokens: Dynamic Grouping with Hierarchical BPE
Rares Dolga, Lucas Maystre, Tudor Berariu +1
Subword tokenization methods like Byte Pair Encoding (BPE) are widely used in large language models due to their balance of vocabulary compactness and representational power. Howev…
Survival Analysis for Idiopathic Pulmonary Fibrosis using CT Images and Incomplete Clinical Data
Ahmed H. Shahin, Joseph Jacob, Daniel C. Alexander +1
Idiopathic Pulmonary Fibrosis (IPF) is an inexorably progressive fibrotic lung disease with a variable and unpredictable rate of progression. CT scans of the lungs inform clinical…
Sample Efficient Model Evaluation
Emine Yilmaz, Peter Hayes, Raza Habib +2
Labelling data is a major practical bottleneck in training and testing classifiers. Given a collection of unlabelled data points, we address how to select which subset to label to…
Gaussian mixture models with Wasserstein distance
Benoit Gaujac, Ilya Feige, David Barber
Generative models with both discrete and continuous latent variables are highly motivated by the structure of many real-world data sets. They present, however, subtleties in traini…
CenTime: Event-Conditional Modelling of Censoring in Survival Analysis
Ahmed H. Shahin, An Zhao, Alexander C. Whitehead +3
Survival analysis is a valuable tool for estimating the time until specific events, such as death or cancer recurrence, based on baseline observations. This is particularly useful…
Parallel Neural Local Lossless Compression
Mingtian Zhang, James Townsend, Ning Kang +1
The recently proposed Neural Local Lossless Compression (NeLLoC), which is based on a local autoregressive model, has achieved state-of-the-art (SOTA) out-of-distribution (OOD) gen…
Active Preference Learning for Large Language Models
William Muldrew, Peter Hayes, Mingtian Zhang +1
As large language models (LLMs) become more capable, fine-tuning techniques for aligning with human intent are increasingly important. A key consideration for aligning these models…
Integrated Weak Learning
Peter Hayes, Mingtian Zhang, Raza Habib +3
We introduce Integrated Weak Learning, a principled framework that integrates weak supervision into the training process of machine learning models. Our approach jointly trains the…
Addressing Catastrophic Forgetting in Few-Shot Problems
Pauching Yap, Hippolyt Ritter, David Barber
Neural networks are known to suffer from catastrophic forgetting when trained on sequential datasets. While there have been numerous attempts to solve this problem in large-scale s…
Practical Gauss-Newton Optimisation for Deep Learning
Aleksandar Botev, Hippolyt Ritter, David Barber
We present an efficient block-diagonal ap- proximation to the Gauss-Newton matrix for feedforward neural networks. Our result- ing algorithm is competitive against state- of-the-ar…
A hybrid CNN-RNN approach for survival analysis in a Lung Cancer Screening study
Yaozhi Lu, Shahab Aslani, An Zhao +5
In this study, we present a hybrid CNN-RNN approach to investigate long-term survival of subjects in a lung cancer screening study. Subjects who died of cardiovascular and respirat…
Practical Lossless Compression with Latent Variables using Bits Back Coding
James Townsend, Tom Bird, David Barber
Deep latent variable models have seen recent success in many data domains. Lossless compression is an application of these models which, despite having the potential to be highly u…
HiLLoC: Lossless Image Compression with Hierarchical Latent Variable Models
James Townsend, Thomas Bird, Julius Kunze +1
We make the following striking observation: fully convolutional VAE models trained on 32x32 ImageNet can generalize well, not just to 64x64 but also to far larger photographs, with…
Learning Deep-Latent Hierarchies by Stacking Wasserstein Autoencoders
Benoit Gaujac, Ilya Feige, David Barber
Probabilistic models with hierarchical-latent-variable structures provide state-of-the-art results amongst non-autoregressive, unsupervised density-based models. However, the most…
Mafin: Enhancing Black-Box Embeddings with Model Augmented Fine-Tuning
Mingtian Zhang, Shawn Lan, Peter Hayes +1
Retrieval Augmented Generation (RAG) has emerged as an effective solution for mitigating hallucinations in Large Language Models (LLMs). The retrieval stage in RAG typically involv…
Stochastic Variational Optimization
Thomas Bird, Julius Kunze, David Barber
Variational Optimization forms a differentiable upper bound on an objective. We show that approaches such as Natural Evolution Strategies and Gaussian Perturbation, are special cas…
Generalized Multiple Intent Conditioned Slot Filling
Harshil Shah, Arthur Wilcke, Marius Cobzarenco +3
Natural language understanding includes the tasks of intent detection (identifying a user's objectives) and slot filling (extracting the entities relevant to those objectives). Pri…
Generalization Gap in Amortized Inference
Mingtian Zhang, Peter Hayes, David Barber
The ability of likelihood-based probabilistic models to generalize to unseen data is central to many machine learning applications such as lossless compression. In this work, we st…
On solving Ordinary Differential Equations using Gaussian Processes
David Barber
We describe a set of Gaussian Process based approaches that can be used to solve non-linear Ordinary Differential Equations. We suggest an explicit probabilistic solver and two imp…
RotRNN: Modelling Long Sequences with Rotations
Kai Biegun, Rares Dolga, Jake Cunningham +1
Linear recurrent neural networks, such as State Space Models (SSMs) and Linear Recurrent Units (LRUs), have recently shown state-of-the-art performance on long sequence modelling b…
Variational f-divergence Minimization
Mingtian Zhang, Thomas Bird, Raza Habib +2
Probabilistic models are often trained by maximum likelihood, which corresponds to minimizing a specific f-divergence between the model and data distribution. In light of recent su…
Representation Learning for High-Dimensional Data Collection under Local Differential Privacy
Alex Mansbridge, Gregory Barbour, Davide Piras +4
The collection of individuals' data has become commonplace in many industries. Local differential privacy (LDP) offers a rigorous approach to preserving privacy whereby the individ…
DiffRatio: Training One-Step Diffusion Models Without Teacher Supervision
Wenlin Chen, Mingtian Zhang, Jiajun He +4
Score-based distillation methods (e.g., variational score distillation) train one-step diffusion models by first pre-training a teacher score model and then distilling it into a on…
Reducing the Computational Cost of Deep Generative Models with Binary Neural Networks
Thomas Bird, Friso H. Kingma, David Barber
Deep generative models provide a powerful set of tools to understand real-world data. But as these models improve, they increase in size and complexity, so their computational cost…
Smoothed Q-learning
David Barber
In Reinforcement Learning the Q-learning algorithm provably converges to the optimal solution. However, as others have demonstrated, Q-learning can also overestimate the values and…
Variational Optimization
Joe Staines, David Barber
We discuss a general technique that can be used to form a differentiable bound on the optima of non-differentiable or discrete objective functions. We form a unified description of…
Gaussian Mean Field Regularizes by Limiting Learned Information
Julius Kunze, Louis Kirsch, Hippolyt Ritter +1
Variational inference with a factorized Gaussian posterior estimate is a widely used approach for learning parameters and hidden variables. Empirically, a regularizing effect can b…
Beyond Internal Data: Constructing Complete Datasets for Fairness Testing
Varsha Ramineni, Hossein A. Rahmani, Emine Yilmaz +1
As AI becomes prevalent in high-risk domains and decision-making, it is essential to test for potential harms and biases. This urgency is reflected by the global emergence of AI re…
Diffusive Gibbs Sampling
Wenlin Chen, Mingtian Zhang, Brooks Paige +2
The inadequate mixing of conventional Markov Chain Monte Carlo (MCMC) methods for multi-modal distributions presents a significant challenge in practical applications such as Bayes…
Locally-Contextual Nonlinear CRFs for Sequence Labeling
Harshil Shah, Tim Xiao, David Barber
Linear chain conditional random fields (CRFs) combined with contextual word embeddings have achieved state of the art performance on sequence labeling tasks. In many of these tasks…
Improving VAE-based Representation Learning
Mingtian Zhang, Tim Z. Xiao, Brooks Paige +1
Latent variable models like the Variational Auto-Encoder (VAE) are commonly used to learn representations of images. However, for downstream tasks like semantic classification, the…
Clique Matrices for Statistical Graph Decomposition and Parameterising Restricted Positive Definite Matrices
David Barber
We introduce Clique Matrices as an alternative representation of undirected graphs, being a generalisation of the incidence matrix representation. Here we use clique matrices to de…
Bayesian Conditional Cointegration
Chris Bracegirdle, David Barber
Cointegration is an important topic for time-series, and describes a relationship between two series in which a linear combination is stationary. Classically, the test for cointegr…
Wider and Deeper, Cheaper and Faster: Tensorized LSTMs for Sequence Learning
Zhen He, Shaobing Gao, Liang Xiao +3
Long Short-Term Memory (LSTM) is a popular approach to boosting the ability of Recurrent Neural Networks to store longer term temporal information. The capacity of an LSTM network…
Unifying Linear-Time Attention via Latent Probabilistic Modelling
Rares Dolga, Lucas Maystre, Marius Cobzarenco +1
Transformers have achieved state-of-the-art results across a range of domains, but their quadratic attention mechanism poses significant challenges for long-sequence modelling. Rec…
Modular Networks: Learning to Decompose Neural Computation
Louis Kirsch, Julius Kunze, David Barber
Scaling model capacity has been vital in the success of deep learning. For a typical network, necessary compute resources and training time grow dramatically with model size. Condi…
Online Structured Laplace Approximations For Overcoming Catastrophic Forgetting
Hippolyt Ritter, Aleksandar Botev, David Barber
We introduce the Kronecker factored online Laplace approximation for overcoming catastrophic forgetting in neural networks. The method is grounded in a Bayesian online learning fra…
Dealing with a large number of classes -- Likelihood, Discrimination or Ranking?
David Barber, Aleksandar Botev
We consider training probabilistic classifiers in the case of a large number of classes. The number of classes is assumed too large to perform exact normalisation over all classes.…
Towards Healing the Blindness of Score Matching
Mingtian Zhang, Oscar Key, Peter Hayes +3
Score-based divergences have been widely used in machine learning and statistics applications. Despite their empirical success, a blindness problem has been observed when using the…
Beyond Internal Data: Bounding and Estimating Fairness from Incomplete Data
Varsha Ramineni, Hossein A. Rahmani, Emine Yilmaz +1
Ensuring fairness in AI systems is critical, especially in high-stakes domains such as lending, hiring, and healthcare. This urgency is reflected in emerging global regulations tha…
Toward Autonomous UI Exploration: The UIExplorer Benchmark
Andrei Cristian Nica, Akshaya Vishnu Kudlu Shanbhogue, Harshil Shah +4
Autonomous agents must know how to explore user interfaces (UIs) for reliable task solving, yet systematic evaluation of this crucial phase is lacking. We introduce UIExplore-Bench…
Moment Matching Denoising Gibbs Sampling
Mingtian Zhang, Alex Hawkins-Hooker, Brooks Paige +1
Energy-Based Models (EBMs) offer a versatile framework for modeling complex data distributions. However, training and sampling from EBMs continue to pose significant challenges. Th…
Thinking Fast and Slow with Deep Learning and Tree Search
Thomas Anthony, Zheng Tian, David Barber
Sequential decision making problems, such as structured prediction, robotic control, and game playing, require a combination of planning policies and generalisation of those plans.…
On the Computational Complexity of Stochastic Controller Optimization in POMDPs
Nikos Vlassis, Michael L. Littman, David Barber
We show that the problem of finding an optimal stochastic 'blind' controller in a Markov decision process is an NP-hard problem. The corresponding decision problem is NP-hard, in P…
Private Machine Learning via Randomised Response
David Barber
We introduce a general learning framework for private machine learning based on randomised response. Our assumption is that all actors are potentially adversarial and as such we tr…
Spread Divergence
Mingtian Zhang, Peter Hayes, Tom Bird +2
For distributions and with different supports or undefined densities, the divergence may not exist. We define a Sprea…
Adaptive Optimization with Examplewise Gradients
Julius Kunze, James Townsend, David Barber
We propose a new, more general approach to the design of stochastic gradient-based optimization methods for machine learning. In this new framework, optimizers assume access to a b…