1 paper
Asma Ghandeharioun, Avi Caciularu, Adam Pearce +2
Understanding the internal representations of large language models (LLMs) can help explain models' behavior and verify their alignment with human values. Given the capabilities of…