1 paper
Chuanhao Yan, Xuhan Huang, Yawen Duan +4
Dense pretrained transformers do not naturally expose interpretable units for circuit extraction. Existing approaches obtain such units by learning auxiliary sparse representations…