6 papers
Deep Learning Models Also Recall Features
Pierre Beckmann
Recent work in mechanistic interpretability has studied how large language models recall facts stored in their weights. This paper argues that factual recall points to something br…
SciR: A Controllable Benchmark for Scientific Reasoning in LLMs
Pierre Beckmann, Marco Valentino, Andre Freitas
Three paradigmatic forms of inference recur across scientific reasoning: deduction, induction, and causal abduction. Reliably evaluating LLMs on these in scientific settings is cur…
Probing Persona-Dependent Preferences in Language Models
Oscar Gilg, Pierre Beckmann, Daniel Paleka +1
Large language models (LLMs) can be said to have preferences: they reliably pick certain tasks and outputs over others, and preferences shaped by post-training and system prompts a…
Where is the Mind? Persona Vectors and LLM Individuation
Pierre Beckmann, Patrick Butlin
The individuation problem for large language models asks which entities associated with them, if any, should be identified as minds. We approach this problem through mechanistic in…
Mechanistic Indicators of Understanding in Large Language Models
Pierre Beckmann, Matthieu Queloz
Large language models (LLMs) are often portrayed as merely imitating linguistic patterns without genuine understanding. We argue that recent findings in mechanistic interpretabilit…
Adaptive LLM-Symbolic Reasoning via Dynamic Logical Solver Composition
Lei Xu, Pierre Beckmann, Marco Valentino +1
Neuro-symbolic NLP methods aim to leverage the complementary strengths of large language models and formal logical solvers. However, current approaches are mostly static in nature,…