3 papers
cs.CL2025
In-Distribution Steering: Balancing Control and Coherence in Language Model Generation
Arthur Vogels, Benjamin Wong, Yann Choho +2
Activation steering methods control large language model (LLM) behavior by modifying internal activations at inference time. However, most existing activation steering methods rely…
cs.CL2025
NeuroFaith: Evaluating LLM Self-Explanation Faithfulness via Internal Representation Alignment
Milan Bhan, Jean-Noel Vittaut, Nicolas Chesneau +2
Large Language Models (LLMs) can generate plausible free text self-explanations to justify their answers. However, these natural language explanations may not accurately reflect th…
cs.CL2025
Towards Achieving Concept Completeness for Textual Concept Bottleneck Models
Milan Bhan, Yann Choho, Pierre Moreau +3
Textual Concept Bottleneck Models (TCBMs) are interpretable-by-design models for text classification that predict a set of salient concepts before making the final prediction. This…