activity
20232026
most citedWhy do explanations fail? A typology and discussion on failures in XAI

4 citations · 6 across the 9 of their papers we have counts for

collaborators

12 papers

cs.AI2026

Diversified and Perceptible Counterfactual Examples Leveraging Expert Knowledge

Akram Bensalem, Fahima Djelil, Marie-Jeanne Lesot +1

CounterFactual Examples (CFEs) are a cornerstone of eXplainable Artificial Intelligence (XAI), offering local, post hoc, and model-agnostic explanations by identifying minimal inpu…

cs.AI2026

Aggregative Semantics for Quantitative Bipolar Argumentation Frameworks

Yann Munro, Isabelle Bloch, Marie-Jeanne Lesot

Formal argumentation is being used increasingly in artificial intelligence as an effective and understandable way to model potentially conflicting pieces of information, called arg…

cs.CL2026

Alignment Reduces Expressed but Not Encoded Gender Bias: A Unified Framework and Study

Nour Bouchouchi, Thibault Laugel, Xavier Renard +3

During training, Large Language Models (LLMs) learn social regularities that can lead to gender bias in downstream applications. Most mitigation efforts focus on reducing bias in g…

cs.AI2025

Metric assessment protocol in the context of answer fluctuation on MCQ tasks

Ekaterina Goliakova, Xavier Renard, Marie-Jeanne Lesot +3

Using multiple-choice questions (MCQs) has become a standard for assessing LLM capabilities efficiently. A variety of metrics can be employed for this task. However, previous resea…

cs.CL2025

NeuroFaith: Evaluating LLM Self-Explanation Faithfulness via Internal Representation Alignment

Milan Bhan, Jean-Noel Vittaut, Nicolas Chesneau +2

Large Language Models (LLMs) can generate plausible free text self-explanations to justify their answers. However, these natural language explanations may not accurately reflect th…

cs.CL2025

Towards Achieving Concept Completeness for Textual Concept Bottleneck Models

Milan Bhan, Yann Choho, Pierre Moreau +3

Textual Concept Bottleneck Models (TCBMs) are interpretable-by-design models for text classification that predict a set of salient concepts before making the final prediction. This…