1 paper · 1 filter
Jungwoo Park, Young Jin Ahn, Kee-Eung Kim +1
Understanding the internal computations of large language models (LLMs) is crucial for aligning them with human values and preventing undesirable behaviors like toxic content gener…