1 citations · 1 across the 12 of their papers we have counts for
1 paper · 1 filter
Miao Yu, Siyuan Fu, Moayad Aloqaily +6
Mechanistic interpretability reveals that safety-critical behaviors (e.g., alignment, jailbreak, backdoor) in Large Language Models (LLMs) are grounded in specialized functional co…