1 citations · 1 across the 8 of their papers we have counts for
3 papers · 1 filter
Limitations of Automated Simulatability: LLM Simulators Can Bypass Explanations
Antonin Poché, Fanny Jourdan, Nils Feldhus +6
Simulatability is an evaluation protocol for explanations that quantifies their usefulness by how well they help a user predict a task model's outputs. Since human evaluation is co…
When Do Concepts Become Functionally Sufficient During Language-Model Training?
Raphael Bernas, Paul G. Chevalier, Fanny Jourdan +1
Understanding a model and its learning mechanisms in depth requires identifying when its internal structures become useful, rather than simply looking at the final state. We study…
Revisiting Anisotropy in Language Transformers: The Geometry of Learning Dynamics
Raphael Bernas, Fanny Jourdan, Antonin Poché +1
Since their introduction, Transformer architectures have dominated Natural Language Processing (NLP). However, recent research has highlighted an inherent anisotropy phenomenon in…