1 paper · 1 filter
Dip Roy, Rajiv Misra, Sanjay Kumar Singh +1
Activation-based linear probing is widely proposed as a method for both detecting and correcting hallucinations in autoregressive language models. We present an empirical study acr…