2 papers
cs.CL2026
Limitations of Automated Simulatability: LLM Simulators Can Bypass Explanations
Antonin Poché, Fanny Jourdan, Nils Feldhus +6
Simulatability is an evaluation protocol for explanations that quantifies their usefulness by how well they help a user predict a task model's outputs. Since human evaluation is co…
cs.CL2023
TaCo: Targeted Concept Erasure Prevents Non-Linear Classifiers From Detecting Protected Attributes
Fanny Jourdan, Louis Béthune, Agustin Picard +2
Ensuring fairness in NLP models is crucial, as they often encode sensitive attributes like gender and ethnicity, leading to biased outcomes. Current concept erasure methods attempt…