1 paper
Giannis Kalyvas, Giorgos Filandrianos, Orfeas Menis Mastromichalakis +2
Large language models can easily generate explanations for their own outputs, but such self-explanations are not necessarily faithful to the model's behavior. We study this issue t…