6 papers
MUtE: A Dual Framework for Concept Erasure and Counterfactual Interventions
Antoine Saillenfest
Erasing concept-specific information from representations has been proven useful for mitigating bias or interpreting model decisions. The joint objective is to transform the origin…
Single-Query Black-Box Calibration Auditing via Logit Bias
Roman Plaud, Antoine Saillenfest, Matthieu Labeau +2
Evaluating the calibration of Large Language Models (LLMs) is critical for their safe deployment as zero-shot classifiers. Yet, commercial API providers increasingly hide the conti…
Tailoring Strictly Proper Scoring Rules for Downstream Tasks: An Application to Causal Inference
Roman Plaud, Alexandre Perez-Lebel, Antoine Saillenfest +4
Probabilistic models are typically trained using task-agnostic objectives like log-loss, which can lead to significant errors in downstream estimation. This disconnect is especiall…
Nonlinear Concept Erasure: a Density Matching Approach
Antoine Saillenfest, Pirmin Lemberger
Ensuring that neural models used in real-world applications cannot infer sensitive information, such as demographic attributes like gender or race, from text representations is a c…
To Each Metric Its Decoding: Post-Hoc Optimal Decision Rules of Probabilistic Hierarchical Classifiers
Roman Plaud, Alexandre Perez-Lebel, Matthieu Labeau +2
Hierarchical classification offers an approach to incorporate the concept of mistake severity by leveraging a structured, labeled hierarchy. However, decoding in such settings freq…
Revisiting Hierarchical Text Classification: Inference and Metrics
Roman Plaud, Matthieu Labeau, Antoine Saillenfest +1
Hierarchical text classification (HTC) is the task of assigning labels to a text within a structured space organized as a hierarchy. Recent works treat HTC as a conventional multil…