1 paper · 1 filter
Stefan Szeider
We present semantic invariance testing, a method to test whether LLM self-explanations are faithful. A faithful self-report should remain stable when only the semantic context chan…