papers

Publications (60)

cs.LG2025

VarDiU: A Variational Diffusive Upper Bound for One-Step Diffusion Distillation

Leyang Wang, Mingtian Zhang, Zijing Ou +1

Recently, diffusion distillation methods have compressed thousand-step teacher diffusion models into one-step student generators while preserving sample quality. Most existing appr…

cs.CL2018

Generating Sentences Using a Dynamic Canvas

Harshil Shah, Bowen Zheng, David Barber

We introduce the Attentive Unsupervised Text (W)riter (AUTR), which is a word level generative model for natural language. It uses a recurrent neural network with a dynamic attenti…

cs.CL2018

Generative Neural Machine Translation

Harshil Shah, David Barber

We introduce Generative Neural Machine Translation (GNMT), a latent variable architecture which is designed to model the semantics of the source and target sentences. We modify an…

cs.CV2019

Tracking by Animation: Unsupervised Learning of Multi-Object Attentive Trackers

Zhen He, Jian Li, Daxue Liu +2

Online Multi-Object Tracking (MOT) from videos is a challenging computer vision task which has been extensively studied for decades. Most of the existing MOT algorithms are based o…

cs.LG2025

Incremental Sequence Classification with Temporal Consistency

Lucas Maystre, Gabriel Barello, Tudor Berariu +5

We address the problem of incremental sequence classification, where predictions are updated as new elements in the sequence are revealed. Drawing on temporal-difference learning f…

stat.ML2020

Learning disentangled representations with the Wasserstein Autoencoder

Benoit Gaujac, Ilya Feige, David Barber

Disentangled representation learning has undoubtedly benefited from objective function surgery. However, a delicate balancing act of tuning is still required in order to trade off…

q-bio.NC2005

Optimal Spike-Timing Dependent Plasticity for Precise Action Potential Firing

Jean-Pascal Pfister, Taro Toyoizumi, David Barber +1

In timing-based neural codes, neurons have to emit action potentials at precise moments in time. We use a supervised learning paradigm to derive a synaptic update rule that optimiz…

cs.LG2025

Training Neural Samplers with Reverse Diffusive KL Divergence

Jiajun He, Wenlin Chen, Mingtian Zhang +2

Training generative models to sample from unnormalized density functions is an important and challenging task in machine learning. Traditional training methods often rely on the re…

stat.ML2018

Improving latent variable descriptiveness with AutoGen

Alex Mansbridge, Roberto Fierimonte, Ilya Feige +1

Powerful generative models, particularly in Natural Language Modelling, are commonly trained by maximizing a variational lower bound on the data log likelihood. These models often…

math.OC2015

An Approximate Newton Method for Markov Decision Processes

Thomas Furmston, David Barber

Gradient-based algorithms are one of the methods of choice for the optimisation of Markov Decision Processes. In this article we will present a novel approximate Newton algorithm f…

eess.SY2012

Efficient Inference in Markov Control Problems

Thomas Furmston, David Barber

Markov control algorithms that perform smooth, non-greedy updates of the policy have been shown to be very general and versatile, with policy gradient and Expectation Maximisation…

stat.ML2016

Nesterov's Accelerated Gradient and Momentum as approximations to Regularised Update Descent

Aleksandar Botev, Guy Lever, David Barber

We present a unifying framework for adapting the update direction in gradient-based iterative optimization methods. As natural special cases we re-derive classical momentum and Nes…

cs.LG2025

Improving Probabilistic Diffusion Models With Optimal Diagonal Covariance Matching

Zijing Ou, Mingtian Zhang, Andi Zhang +3

The probabilistic diffusion model has become highly effective across various domains. Typically, sampling from a diffusion model involves using a denoising distribution characteriz…

cs.CL2025

From Characters to Tokens: Dynamic Grouping with Hierarchical BPE

Rares Dolga, Lucas Maystre, Tudor Berariu +1

Subword tokenization methods like Byte Pair Encoding (BPE) are widely used in large language models due to their balance of vocabulary compactness and representational power. Howev…

eess.IV2022

Survival Analysis for Idiopathic Pulmonary Fibrosis using CT Images and Incomplete Clinical Data

Ahmed H. Shahin, Joseph Jacob, Daniel C. Alexander +1

Idiopathic Pulmonary Fibrosis (IPF) is an inexorably progressive fibrotic lung disease with a variable and unpredictable rate of progression. CT scans of the lungs inform clinical…

cs.LG2021

Sample Efficient Model Evaluation

Emine Yilmaz, Peter Hayes, Raza Habib +2

Labelling data is a major practical bottleneck in training and testing classifiers. Given a collection of unlabelled data points, we address how to select which subset to label to…

stat.ML2018

Gaussian mixture models with Wasserstein distance

Benoit Gaujac, Ilya Feige, David Barber

Generative models with both discrete and continuous latent variables are highly motivated by the structure of many real-world data sets. They present, however, subtleties in traini…

cs.LG2024

CenTime: Event-Conditional Modelling of Censoring in Survival Analysis

Ahmed H. Shahin, An Zhao, Alexander C. Whitehead +3

Survival analysis is a valuable tool for estimating the time until specific events, such as death or cancer recurrence, based on baseline observations. This is particularly useful…

eess.IV2022

Parallel Neural Local Lossless Compression

Mingtian Zhang, James Townsend, Ning Kang +1

The recently proposed Neural Local Lossless Compression (NeLLoC), which is based on a local autoregressive model, has achieved state-of-the-art (SOTA) out-of-distribution (OOD) gen…

cs.LG2024

Active Preference Learning for Large Language Models

William Muldrew, Peter Hayes, Mingtian Zhang +1

As large language models (LLMs) become more capable, fine-tuning techniques for aligning with human intent are increasingly important. A key consideration for aligning these models…

cs.LG2022

Integrated Weak Learning

Peter Hayes, Mingtian Zhang, Raza Habib +3

We introduce Integrated Weak Learning, a principled framework that integrates weak supervision into the training process of machine learning models. Our approach jointly trains the…

cs.LG2021

Addressing Catastrophic Forgetting in Few-Shot Problems

Pauching Yap, Hippolyt Ritter, David Barber

Neural networks are known to suffer from catastrophic forgetting when trained on sequential datasets. While there have been numerous attempts to solve this problem in large-scale s…

stat.ML2017

Practical Gauss-Newton Optimisation for Deep Learning

Aleksandar Botev, Hippolyt Ritter, David Barber

We present an efficient block-diagonal ap- proximation to the Gauss-Newton matrix for feedforward neural networks. Our result- ing algorithm is competitive against state- of-the-ar…

cs.LG2023

A hybrid CNN-RNN approach for survival analysis in a Lung Cancer Screening study

Yaozhi Lu, Shahab Aslani, An Zhao +5

In this study, we present a hybrid CNN-RNN approach to investigate long-term survival of subjects in a lung cancer screening study. Subjects who died of cardiovascular and respirat…

cs.LG2019

Practical Lossless Compression with Latent Variables using Bits Back Coding

James Townsend, Tom Bird, David Barber

Deep latent variable models have seen recent success in many data domains. Lossless compression is an application of these models which, despite having the potential to be highly u…

eess.IV2019

HiLLoC: Lossless Image Compression with Hierarchical Latent Variable Models

James Townsend, Thomas Bird, Julius Kunze +1

We make the following striking observation: fully convolutional VAE models trained on 32x32 ImageNet can generalize well, not just to 64x64 but also to far larger photographs, with…

stat.ML2020

Learning Deep-Latent Hierarchies by Stacking Wasserstein Autoencoders

Benoit Gaujac, Ilya Feige, David Barber

Probabilistic models with hierarchical-latent-variable structures provide state-of-the-art results amongst non-autoregressive, unsupervised density-based models. However, the most…

cs.LG2024

Mafin: Enhancing Black-Box Embeddings with Model Augmented Fine-Tuning

Mingtian Zhang, Shawn Lan, Peter Hayes +1

Retrieval Augmented Generation (RAG) has emerged as an effective solution for mitigating hallucinations in Large Language Models (LLMs). The retrieval stage in RAG typically involv…

stat.ML2018

Stochastic Variational Optimization

Thomas Bird, Julius Kunze, David Barber

Variational Optimization forms a differentiable upper bound on an objective. We show that approaches such as Natural Evolution Strategies and Gaussian Perturbation, are special cas…

cs.CL2023

Generalized Multiple Intent Conditioned Slot Filling

Harshil Shah, Arthur Wilcke, Marius Cobzarenco +3

Natural language understanding includes the tasks of intent detection (identifying a user's objectives) and slot filling (extracting the entities relevant to those objectives). Pri…

stat.ML2022

Generalization Gap in Amortized Inference

Mingtian Zhang, Peter Hayes, David Barber

The ability of likelihood-based probabilistic models to generalize to unseen data is central to many machine learning applications such as lossless compression. In this work, we st…

stat.ME2014

On solving Ordinary Differential Equations using Gaussian Processes

David Barber

We describe a set of Gaussian Process based approaches that can be used to solve non-linear Ordinary Differential Equations. We suggest an explicit probabilistic solver and two imp…

cs.LG2026

RotRNN: Modelling Long Sequences with Rotations

Kai Biegun, Rares Dolga, Jake Cunningham +1

Linear recurrent neural networks, such as State Space Models (SSMs) and Linear Recurrent Units (LRUs), have recently shown state-of-the-art performance on long sequence modelling b…

stat.ML2024

Variational f-divergence Minimization

Mingtian Zhang, Thomas Bird, Raza Habib +2

Probabilistic models are often trained by maximum likelihood, which corresponds to minimizing a specific f-divergence between the model and data distribution. In light of recent su…

cs.LG2022

Representation Learning for High-Dimensional Data Collection under Local Differential Privacy

Alex Mansbridge, Gregory Barbour, Davide Piras +4

The collection of individuals' data has become commonplace in many industries. Local differential privacy (LDP) offers a rigorous approach to preserving privacy whereby the individ…

cs.LG2026

DiffRatio: Training One-Step Diffusion Models Without Teacher Supervision

Wenlin Chen, Mingtian Zhang, Jiajun He +4

Score-based distillation methods (e.g., variational score distillation) train one-step diffusion models by first pre-training a teacher score model and then distilling it into a on…

cs.LG2021

Reducing the Computational Cost of Deep Generative Models with Binary Neural Networks

Thomas Bird, Friso H. Kingma, David Barber

Deep generative models provide a powerful set of tools to understand real-world data. But as these models improve, they increase in size and complexity, so their computational cost…

cs.LG2023

Smoothed Q-learning

David Barber

In Reinforcement Learning the Q-learning algorithm provably converges to the optimal solution. However, as others have demonstrated, Q-learning can also overestimate the values and…

stat.ML2012

Variational Optimization

Joe Staines, David Barber

We discuss a general technique that can be used to form a differentiable bound on the optima of non-differentiable or discrete objective functions. We form a unified description of…

cs.LG2019

Gaussian Mean Field Regularizes by Limiting Learned Information

Julius Kunze, Louis Kirsch, Hippolyt Ritter +1

Variational inference with a factorized Gaussian posterior estimate is a widely used approach for learning parameters and hidden variables. Empirically, a regularizing effect can b…

cs.LG2025

Beyond Internal Data: Constructing Complete Datasets for Fairness Testing

Varsha Ramineni, Hossein A. Rahmani, Emine Yilmaz +1

As AI becomes prevalent in high-risk domains and decision-making, it is essential to test for potential harms and biases. This urgency is reflected by the global emergence of AI re…

stat.ML2024

Diffusive Gibbs Sampling

Wenlin Chen, Mingtian Zhang, Brooks Paige +2

The inadequate mixing of conventional Markov Chain Monte Carlo (MCMC) methods for multi-modal distributions presents a significant challenge in practical applications such as Bayes…

cs.CL2021

Locally-Contextual Nonlinear CRFs for Sequence Labeling

Harshil Shah, Tim Xiao, David Barber

Linear chain conditional random fields (CRFs) combined with contextual word embeddings have achieved state of the art performance on sequence labeling tasks. In many of these tasks…

stat.ML2022

Improving VAE-based Representation Learning

Mingtian Zhang, Tim Z. Xiao, Brooks Paige +1

Latent variable models like the Variational Auto-Encoder (VAE) are commonly used to learn representations of images. However, for downstream tasks like semantic classification, the…

cs.DM2012

Clique Matrices for Statistical Graph Decomposition and Parameterising Restricted Positive Definite Matrices

David Barber

We introduce Clique Matrices as an alternative representation of undirected graphs, being a generalisation of the incidence matrix representation. Here we use clique matrices to de…

cs.CE2012

Bayesian Conditional Cointegration

Chris Bracegirdle, David Barber

Cointegration is an important topic for time-series, and describes a relationship between two series in which a linear combination is stationary. Classically, the test for cointegr…

stat.ML2017

Wider and Deeper, Cheaper and Faster: Tensorized LSTMs for Sequence Learning

Zhen He, Shaobing Gao, Liang Xiao +3

Long Short-Term Memory (LSTM) is a popular approach to boosting the ability of Recurrent Neural Networks to store longer term temporal information. The capacity of an LSTM network…

cs.CL2025

Unifying Linear-Time Attention via Latent Probabilistic Modelling

Rares Dolga, Lucas Maystre, Marius Cobzarenco +1

Transformers have achieved state-of-the-art results across a range of domains, but their quadratic attention mechanism poses significant challenges for long-sequence modelling. Rec…

cs.LG2018

Modular Networks: Learning to Decompose Neural Computation

Louis Kirsch, Julius Kunze, David Barber

Scaling model capacity has been vital in the success of deep learning. For a typical network, necessary compute resources and training time grow dramatically with model size. Condi…

stat.ML2018

Online Structured Laplace Approximations For Overcoming Catastrophic Forgetting

Hippolyt Ritter, Aleksandar Botev, David Barber

We introduce the Kronecker factored online Laplace approximation for overcoming catastrophic forgetting in neural networks. The method is grounded in a Bayesian online learning fra…

stat.ML2016

Dealing with a large number of classes -- Likelihood, Discrimination or Ranking?

David Barber, Aleksandar Botev

We consider training probabilistic classifiers in the case of a large number of classes. The number of classes is assumed too large to perform exact normalisation over all classes.…

stat.ML2025

Towards Healing the Blindness of Score Matching

Mingtian Zhang, Oscar Key, Peter Hayes +3

Score-based divergences have been widely used in machine learning and statistics applications. Despite their empirical success, a blindness problem has been observed when using the…

cs.LG2025

Beyond Internal Data: Bounding and Estimating Fairness from Incomplete Data

Varsha Ramineni, Hossein A. Rahmani, Emine Yilmaz +1

Ensuring fairness in AI systems is critical, especially in high-stakes domains such as lending, hiring, and healthcare. This urgency is reflected in emerging global regulations tha…

cs.LG2025

Toward Autonomous UI Exploration: The UIExplorer Benchmark

Andrei Cristian Nica, Akshaya Vishnu Kudlu Shanbhogue, Harshil Shah +4

Autonomous agents must know how to explore user interfaces (UIs) for reliable task solving, yet systematic evaluation of this crucial phase is lacking. We introduce UIExplore-Bench…

stat.ML2024

Moment Matching Denoising Gibbs Sampling

Mingtian Zhang, Alex Hawkins-Hooker, Brooks Paige +1

Energy-Based Models (EBMs) offer a versatile framework for modeling complex data distributions. However, training and sampling from EBMs continue to pose significant challenges. Th…

cs.AI2017

Thinking Fast and Slow with Deep Learning and Tree Search

Thomas Anthony, Zheng Tian, David Barber

Sequential decision making problems, such as structured prediction, robotic control, and game playing, require a combination of planning policies and generalisation of those plans.…

cs.CC2012

On the Computational Complexity of Stochastic Controller Optimization in POMDPs

Nikos Vlassis, Michael L. Littman, David Barber

We show that the problem of finding an optimal stochastic 'blind' controller in a Markov decision process is an NP-hard problem. The corresponding decision problem is NP-hard, in P…

cs.LG2020

Private Machine Learning via Randomised Response

David Barber

We introduce a general learning framework for private machine learning based on randomised response. Our assumption is that all actors are potentially adversarial and as such we tr…

stat.ML2022

Spread Divergence

Mingtian Zhang, Peter Hayes, Tom Bird +2

For distributions and with different supports or undefined densities, the divergence may not exist. We define a Sprea…

cs.LG2021

Adaptive Optimization with Examplewise Gradients

Julius Kunze, James Townsend, David Barber

We propose a new, more general approach to the design of stochastic gradient-based optimization methods for machine learning. In this new framework, optimizers assume access to a b…