1 paper · 1 filter
Miriam Rateike, Celia Cintas, John Wamburu +2
We propose an auditing method to identify whether a large language model (LLM) encodes patterns such as hallucinations in its internal states, which may propagate to downstream tas…