39 citations · 155 across the 12 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2024★ 2 cited
Sabotage Evaluations for Frontier Models
Joe Benton, Misha Wagner, Eric Christiansen +13
Sufficiently capable models could subvert human oversight and decision-making in important contexts. For example, in the context of AI development, models could covertly sabotage e…
cs.LG2022★ 5 cited
Engineering Monosemanticity in Toy Models
Adam S. Jermyn, Nicholas Schiefer, Evan Hubinger
In some neural networks, individual neurons correspond to natural ``features'' in the input. Such \emph{monosemantic} neurons are of great help in interpretability studies, as they…