activity
20242026
collaborators

10 papers

cs.CL2026

Alignment Reduces Expressed but Not Encoded Gender Bias: A Unified Framework and Study

Nour Bouchouchi, Thibault Laugel, Xavier Renard +3

During training, Large Language Models (LLMs) learn social regularities that can lead to gender bias in downstream applications. Most mitigation efforts focus on reducing bias in g…

cs.AI2026

Aggregative Semantics for Quantitative Bipolar Argumentation Frameworks

Yann Munro, Isabelle Bloch, Marie-Jeanne Lesot

Formal argumentation is being used increasingly in artificial intelligence as an effective and understandable way to model potentially conflicting pieces of information, called arg…

cs.CL2026

NeuroFaith: Evaluating LLM Self-Explanation Faithfulness via Internal Representation Alignment

Milan Bhan, Jean-Noel Vittaut, Nicolas Chesneau +2

Large Language Models (LLMs) can generate plausible free text self-explanations to justify their answers. However, these natural language explanations may not accurately reflect th…

cs.LG2025

Why do explanations fail? A typology and discussion on failures in XAI

Clara Bove, Thibault Laugel, Marie-Jeanne Lesot +2

As Machine Learning models achieve unprecedented levels of performance, the XAI domain aims at making these models understandable by presenting end-users with intelligible explanat…

cs.AI2025

Metric assessment protocol in the context of answer fluctuation on MCQ tasks

Ekaterina Goliakova, Xavier Renard, Marie-Jeanne Lesot +3

Using multiple-choice questions (MCQs) has become a standard for assessing LLM capabilities efficiently. A variety of metrics can be employed for this task. However, previous resea…

cs.CL2025

Towards Achieving Concept Completeness for Textual Concept Bottleneck Models

Milan Bhan, Yann Choho, Pierre Moreau +3

Textual Concept Bottleneck Models (TCBMs) are interpretable-by-design models for text classification that predict a set of salient concepts before making the final prediction. This…