4 papers · 1 filter
Evaluating the Relevance of Uncertainty Estimators for LLM Hallucination
Yedidia Agnimo, Anna Korba, Annabelle Blangero +2
Large language models (LLMs) are prone to hallucinations, i.e., statements unsupported by the input or training data, hindering reliable deployment. In parallel, numerous uncertain…
NeuroFaith: Evaluating LLM Self-Explanation Faithfulness via Internal Representation Alignment
Milan Bhan, Jean-Noel Vittaut, Nicolas Chesneau +2
Large Language Models (LLMs) can generate plausible free text self-explanations to justify their answers. However, these natural language explanations may not accurately reflect th…
Towards Achieving Concept Completeness for Textual Concept Bottleneck Models
Milan Bhan, Yann Choho, Pierre Moreau +3
Textual Concept Bottleneck Models (TCBMs) are interpretable-by-design models for text classification that predict a set of salient concepts before making the final prediction. This…
Mitigating Text Toxicity with Counterfactual Generation
Milan Bhan, Jean-Noel Vittaut, Nina Achache +5
Toxicity mitigation consists in rephrasing text in order to remove offensive or harmful meaning. Neural natural language processing (NLP) models have been widely used to target and…