1 citations · 1 across the 1 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
MAxBench: A Multinomial Concept Recovery Benchmark
Divya Appapogu, Freya Behrens, Yonatan Belinkov +1
Fine-grained control of language model behaviors (e.g., steering) is among the more actionable outcomes of interpretability research. For binary concepts such as refusal, a single…
cs.LG2023★ 1 cited
Unveiling the Hessian's Connection to the Decision Boundary
Mahalakshmi Sabanayagam, Freya Behrens, Urte Adomaityte +1
Understanding the properties of well-generalizing minima is at the heart of deep learning research. On the one hand, the generalization of neural networks has been connected to the…