papers

Publications (11)

cs.LG2023

Hierarchical Partitioning Forecaster

Christopher Mattern

In this work we consider a new family of algorithms for sequential prediction, Hierarchical Partitioning Forecasters (HPFs). Our goal is to provide appealing theoretical - regret g…

cs.LG2020

Gated Linear Networks

Joel Veness, Tor Lattimore, David Budden +8

This paper presents a new family of backpropagation-free neural architectures, Gated Linear Networks (GLNs). What distinguishes GLNs from contemporary neural networks is the distri…

cs.LG2024

Learning Universal Predictors

Jordi Grau-Moya, Tim Genewein, Marcus Hutter +8

Meta-learning has emerged as a powerful approach to train neural networks to learn new tasks quickly from limited data. Broad exposure to different tasks leads to versatile represe…

cs.IT2018

Generalized Probability Smoothing

Christopher Mattern

In this work we consider a generalized version of Probability Smoothing, the core elementary model for sequential prediction in the state of the art PAQ family of data compression…

cs.LG2017

Online Learning with Gated Linear Networks

Joel Veness, Tor Lattimore, Avishkar Bhoopchand +3

This paper describes a family of probabilistic architectures designed for online learning under the logarithmic loss. Rather than relying on non-linear transfer functions, our meth…

cs.IT2015

On Probability Estimation via Relative Frequencies and Discount

Christopher Mattern

Probability estimation is an elementary building block of every statistical data compression algorithm. In practice probability estimation is often based on relative letter frequen…

cs.IT2013

Mixing Strategies in Data Compression

Christopher Mattern

We propose geometric weighting as a novel method to combine multiple models in data compression. Our results reveal the rationale behind PAQ-weighting and generalize it to a non-bi…

cs.IT2015

On Probability Estimation by Exponential Smoothing

Christopher Mattern

Probability estimation is essential for every statistical data compression algorithm. In practice probability estimation should be adaptive, recent observations should receive a hi…

cs.IT2013

Combining non-stationary prediction, optimization and mixing for data compression

Christopher Mattern

In this paper an approach to modelling nonstationary binary sequences, i.e., predicting the probability of upcoming symbols, is presented. After studying the prediction model we ev…

cs.LG2024

Language Modeling Is Compression

Grégoire Delétang, Anian Ruoss, Paul-Ambroise Duquenne +9

It has long been established that predictive models can be transformed into lossless compressors and vice versa. Incidentally, in recent years, the machine learning community has f…

cs.IT2013

Linear and Geometric Mixtures - Analysis

Christopher Mattern

Linear and geometric mixtures are two methods to combine arbitrary models in data compression. Geometric mixtures generalize the empirically well-performing PAQ7 mixture. Both mixt…