3 papers
cs.CV2026
Rodent-Bench
Thomas Heap, Laurence Aitchison, Emma Cahill +1
We present Rodent-Bench, a novel benchmark designed to evaluate the ability of Multimodal Large Language Models (MLLMs) to annotate rodent behaviour footage. We evaluate state-of-t…
cs.LG2026
Automated Interpretability Metrics Do Not Distinguish Trained and Random Transformers
Thomas Heap, Tim Lawson, Lucy Farnik +1
Sparse autoencoders (SAEs) are widely used to extract sparse, interpretable latents from transformer activations. We test whether commonly used SAE quality metrics and automatic ex…
stat.ML2025
Massively Parallel Expectation Maximization For Approximate Posteriors
Thomas Heap, Sam Bowyer, Laurence Aitchison
Bayesian inference for hierarchical models can be very challenging. MCMC methods have difficulty scaling to large models with many observations and latent variables. While variatio…