activity
20212024
most citedNeural Networks and the Chomsky Hierarchy

45 citations · 93 across the 10 of their papers we have counts for

collaborators
Showing cs.LGShow all

9 papers · 1 filter

cs.LG2024★ 8 cited

Amortized Planning with Large-Scale Transformers: A Case Study on Chess

Anian Ruoss, Grégoire Delétang, Sourabh Medapati +7

This paper uses chess, a landmark planning problem in AI, to assess transformers' performance on a planning task where memorization is futile $\unicode{x2013}$ even at a large scal…

cs.LG2024★ 1 cited

Learning Universal Predictors

Jordi Grau-Moya, Tim Genewein, Marcus Hutter +8

Meta-learning has emerged as a powerful approach to train neural networks to learn new tasks quickly from limited data. Broad exposure to different tasks leads to versatile represe…

cs.LG2023★ 28 cited

Language Modeling Is Compression

Grégoire Delétang, Anian Ruoss, Paul-Ambroise Duquenne +9

It has long been established that predictive models can be transformed into lossless compressors and vice versa. Incidentally, in recent years, the machine learning community has f…

cs.LG2023★ 1 cited

Randomized Positional Encodings Boost Length Generalization of Transformers

Anian Ruoss, Grégoire Delétang, Tim Genewein +5

Transformers have impressive generalization capabilities on tasks with a fixed context length. However, they fail to generalize to sequences of arbitrary length, even for seemingly…

cs.LG2023★ 1 cited

Memory-Based Meta-Learning on Non-Stationary Distributions

Tim Genewein, Grégoire Delétang, Anian Ruoss +7

Memory-based meta-learning is a technique for approximating Bayes-optimal predictors. Under fairly general conditions, minimizing sequential prediction error, measured by the log l…

cs.LG2022★ 45 cited

Neural Networks and the Chomsky Hierarchy

Grégoire Delétang, Anian Ruoss, Jordi Grau-Moya +8

Reliable generalization lies at the heart of safe ML and AI. However, understanding when and how neural networks generalize remains one of the most important unsolved problems in t…