3 papers
cs.CV2025
I Spy With My Model's Eye: Visual Search as a Behavioural Test for MLLMs
John Burden, Jonathan Prunty, Ben Slater +3
Multimodal large language models (MLLMs) achieve strong performance on vision-language tasks, yet their visual processing is opaque. Most black-box evaluations measure task accurac…
cs.CL2025
Mechanistic Decomposition of Sentence Representations
Matthieu Tehenan, Vikram Natarajan, Jonathan Michala +2
Sentence embeddings are central to modern NLP and AI systems, yet little is known about their internal structure. While we can compare these embeddings using measures such as cosin…
cs.AI2025
Linear Spatial World Models Emerge in Large Language Models
Matthieu Tehenan, Christian Bolivar Moya, Tenghai Long +1
Large language models (LLMs) have demonstrated emergent abilities across diverse tasks, raising the question of whether they acquire internal world models. In this work, we investi…