13 papers
HALLMARK: Diagnosing Three Failure Modes in LLM Citation Verifiers
Patrik Reizinger, Wieland Brendel
Large language models (LLMs) now routinely draft literature reviews and assist with academic writing, which means a higher risk of fabricated references: GPTZero found 53 papers wi…
Stabilizing Extrapolation in Looped Transformers via Learned Stochastic Stopping
Hsun-Yu Kuo, El Mahdi Chayti, Patrik Reizinger +2
Looped Transformers, which repeatedly apply a shared transformer block, are an architecturally natural fit for variable-length algorithmic tasks. Although they can exhibit strong l…
From Isolation to Entanglement: When Do Interpretability Methods Identify and Disentangle Known Concepts?
Aaron Mueller, Andrew Lee, Shruti Joshi +3
A goal of interpretability is to recover disentangled representations of latent concepts (features) from the activations of neural networks. The quality of features is typically ev…
Causality is Key for Interpretability Claims to Generalise
Shruti Joshi, Aaron Mueller, David Klindt +3
Interpretability research on large language models (LLMs) has yielded important insights into model behaviour, yet recurring pitfalls persist: findings that do not generalise, and…
Who Guards the Guardians? The Challenges of Evaluating Identifiability of Learned Representations
Shruti Joshi, Théo Saulus, Wieland Brendel +3
Identifiability in representation learning is commonly evaluated using standard metrics (e.g., MCC, DCI, R^2) on synthetic benchmarks with known ground-truth factors. These metrics…
Estimating Treatment Effects with Independent Component Analysis
Patrik Reizinger, Lester Mackey, Wieland Brendel +1
Independent Component Analysis (ICA) uses a measure of non-Gaussianity to identify latent sources from data and estimate their mixing coefficients (Shimizu et al., 2006). Meanwhile…