1 paper
Miao Yu, Siyuan Fu, Moayad Aloqaily +6
Mechanistic interpretability reveals that safety-critical behaviors (e.g., alignment, jailbreak, backdoor) in Large Language Models (LLMs) are grounded in specialized functional co…