2 papers
cs.LG2026
On the Relationship Between Activation Outliers and Feature Death in Sparse Autoencoders
Elana Simon, Etowah Adams, James Zou
Sparse autoencoders (SAEs) decompose neural network activations into interpretable features, but many learned features never activate, a problem called feature death that wastes di…
cs.LG2026
Beyond Linear Steering: Unified Multi-Attribute Control for Language Models
Narmeen Oozeer, Luke Marks, Shreyans Jain +2
Controlling multiple behavioral attributes in large language models (LLMs) at inference time is a challenging problem due to interference between attributes and the limitations of…