Publications (19)
Position: An Empirically Grounded Identifiability Theory Will Accelerate Self-Supervised Learning Research
Patrik Reizinger, Randall Balestriero, David Klindt +1
Self-Supervised Learning (SSL) powers many current AI systems. As research interest and investment grow, the SSL design space continues to expand. The Platonic view of SSL, followi…
Attention-based Curiosity-driven Exploration in Deep Reinforcement Learning
Patrik Reizinger, Márton Szemenyei
Reinforcement Learning enables to train an agent via interaction with the environment. However, in the majority of real-world scenarios, the extrinsic feedback is sparse or not suf…
From Isolation to Entanglement: When Do Interpretability Methods Identify and Disentangle Known Concepts?
Aaron Mueller, Andrew Lee, Shruti Joshi +3
A goal of interpretability is to recover disentangled representations of latent concepts (features) from the activations of neural networks. The quality of features is typically ev…
Stabilizing Extrapolation in Looped Transformers via Learned Stochastic Stopping
Hsun-Yu Kuo, El Mahdi Chayti, Patrik Reizinger +2
Looped Transformers, which repeatedly apply a shared transformer block, are an architecturally natural fit for variable-length algorithmic tasks. Although they can exhibit strong l…
Who Guards the Guardians? The Challenges of Evaluating Identifiability of Learned Representations
Shruti Joshi, Théo Saulus, Wieland Brendel +3
Identifiability in representation learning is commonly evaluated using standard metrics (e.g., MCC, DCI, R^2) on synthetic benchmarks with known ground-truth factors. These metrics…
An Interventional Perspective on Identifiability in Gaussian LTI Systems with Independent Component Analysis
Goutham Rajendran, Patrik Reizinger, Wieland Brendel +1
We investigate the relationship between system identification and intervention design in dynamical systems. While previous research demonstrated how identifiable representation lea…
Causality is Key for Interpretability Claims to Generalise
Shruti Joshi, Aaron Mueller, David Klindt +3
Interpretability research on large language models (LLMs) has yielded important insights into model behaviour, yet recurring pitfalls persist: findings that do not generalise, and…
Estimating Treatment Effects with Independent Component Analysis
Patrik Reizinger, Lester Mackey, Wieland Brendel +1
Independent Component Analysis (ICA) uses a measure of non-Gaussianity to identify latent sources from data and estimate their mixing coefficients (Shimizu et al., 2006). Meanwhile…
Skill Learning via Policy Diversity Yields Identifiable Representations for Reinforcement Learning
Patrik Reizinger, Bálint Mucsányi, Siyuan Guo +3
Self-supervised feature learning and pretraining methods in reinforcement learning (RL) often rely on information-theoretic principles, termed mutual information skill learning (MI…
InfoNCE: Identifying the Gap Between Theory and Practice
Evgenia Rusak, Patrik Reizinger, Attila Juhos +3
Prior theory work on Contrastive Learning via the InfoNCE loss showed that, under certain assumptions, the learned representations recover the ground-truth latent factors. We argue…
Position: Understanding LLMs Requires More Than Statistical Generalization
Patrik Reizinger, Szilvia Ujváry, Anna Mészáros +3
The last decade has seen blossoming research in deep learning theory attempting to answer, "Why does deep learning generalize?" A powerful shift in perspective precipitated this pr…
Cross-Entropy Is All You Need To Invert the Data Generating Process
Patrik Reizinger, Alice Bizeul, Attila Juhos +4
Supervised learning has become a cornerstone of modern machine learning, yet a comprehensive theory explaining its effectiveness remains elusive. Empirical phenomena, such as neura…
From superposition to sparse codes: interpretable representations in neural networks
David Klindt, Charles O'Neill, Patrik Reizinger +2
Understanding how information is represented in neural networks is a fundamental challenge in both neuroscience and artificial intelligence. Despite their nonlinear architectures,…
Identifiable Exchangeable Mechanisms for Causal Structure and Representation Learning
Patrik Reizinger, Siyuan Guo, Ferenc Huszár +2
Identifying latent representations or causal structures is important for good generalization and downstream task performance. However, both fields have been developed rather indepe…
HALLMARK: Diagnosing Three Failure Modes in LLM Citation Verifiers
Patrik Reizinger, Wieland Brendel
Large language models (LLMs) now routinely draft literature reviews and assist with academic writing, which means a higher risk of fabricated references: GPTZero found 53 papers wi…
Out-of-distribution Tests Reveal Compositionality in Chess Transformers
Anna Mészáros, Patrik Reizinger, Ferenc Huszár
Chess is a canonical example of a task that requires rigorous reasoning and long-term planning. Modern decision Transformers - trained similarly to LLMs - are able to learn compete…
Embrace the Gap: VAEs Perform Independent Mechanism Analysis
Patrik Reizinger, Luigi Gresele, Jack Brady +6
Variational autoencoders (VAEs) are a popular framework for modeling complex data distributions; they can be efficiently trained via variational inference by maximizing the evidenc…
Stochastic Weight Matrix-based Regularization Methods for Deep Neural Networks
Patrik Reizinger, Bálint Gyires-Tóth
The aim of this paper is to introduce two widely applicable regularization methods based on the direct modification of weight matrices. The first method, Weight Reinitialization, u…
Rule Extrapolation in Language Models: A Study of Compositional Generalization on OOD Prompts
Anna Mészáros, Szilvia Ujváry, Wieland Brendel +2
LLMs show remarkable emergent abilities, such as inferring concepts from presumably out-of-distribution prompts, known as in-context learning. Though this success is often attribut…