4 papers
Learning Concept Bottleneck Models from Mechanistic Explanations
Antonio De Santis, Schrasing Tong, Marco Brambilla +1
Concept Bottleneck Models (CBMs) aim for ante-hoc interpretability by learning a bottleneck layer that predicts interpretable concepts before the decision. State-of-the-art approac…
Mitigating Bias in Concept Bottleneck Models for Fair and Interpretable Image Classification
Schrasing Tong, Antoine Salaun, Vincent Yuan +2
Ensuring fairness in image classification prevents models from perpetuating and amplifying bias. Concept bottleneck models (CBMs) map images to high-level, human-interpretable conc…
Measuring Perceptions of Fairness in AI Systems: The Effects of Infra-marginality
Schrasing Tong, Minseok Jung, Ilaria Liccardi +1
Differences in data distributions between demographic groups, known as the problem of infra-marginality, complicate how people evaluate fairness in machine learning models. We pres…
Towards Resource Efficient and Interpretable Bias Mitigation in Large Language Models
Schrasing Tong, Eliott Zemour, Jessica Lu +2
Although large language models (LLMs) have demonstrated their effectiveness in a wide range of applications, they have also been observed to perpetuate unwanted biases present in t…