43 citations · 55 across the 8 of their papers we have counts for
1 paper · 1 filter
Asma Ghandeharioun, Avi Caciularu, Adam Pearce +2
Understanding the internal representations of large language models (LLMs) can help explain models' behavior and verify their alignment with human values. Given the capabilities of…