1 paper · 1 filter
Shivam Adarsh, Maria Maistro, Christina Lioma
Large Language Models (LLMs) often encode whether a statement is true as a vector in their residual stream activations. These vectors, also known as truth vectors, have been studie…