5 citations · 5 across the 2 of their papers we have counts for
2 papers
cs.LG2024★ 5 cited
BatchTopK Sparse Autoencoders
Bart Bussmann, Patrick Leask, Neel Nanda
Sparse autoencoders (SAEs) have emerged as a powerful tool for interpreting language model activations by decomposing them into sparse, interpretable features. A popular approach i…
cs.AI2023
CoinRun: Solving Goal Misgeneralisation
Stuart Armstrong, Alexandre Maranhão, Oliver Daniels-Koch +2
Goal misgeneralisation is a key challenge in AI alignment -- the task of getting powerful Artificial Intelligences to align their goals with human intentions and human morality. In…