4 citations · 8 across the 18 of their papers we have counts for
5 papers · 1 filter
Are Single-Token Sparse Autoencoder Features Causally Necessary? Layer-Depth and SAE-Family Effects
Seonglae Cho, Zekun Wu, Kleyton Da Costa +3
Sparse autoencoder (SAE) features are used to interpret and steer large language models, yet nobody has tested whether a feature's causal role is stable across SAE families. Single…
Which Decisions Low-Bit Quantization Breaks, and How to Predict Them
Zekun Wu, Swati Dhiman, Adriano Koshiyama
Quantization is known to hurt below four bits, but nobody can say which of a model's decisions will change at a given bit-width. This matters most where a model acts rather than an…
Towards The Implicit Bias on Multiclass Separable Data Under Norm Constraints
Shengping Xie, Zekun Wu, Quan Chen +1
Implicit bias induced by gradient-based algorithms is essential to the generalization of overparameterized models, yet its mechanisms can be subtle. This work leverages the Normali…
Control Reinforcement Learning: Interpretable Token-Level Steering of LLMs via Sparse Autoencoder Features
Seonglae Cho, Zekun Wu, Adriano Koshiyama
Sparse autoencoders (SAEs) decompose language model activations into interpretable features, but existing methods reveal only which features activate, not which change model output…
The Confidence Manifold: Geometric Structure of Correctness Representations in Language Models
Seonglae Cho, Zekun Wu, Kleyton Da Costa +1
When a language model asserts that "the capital of Australia is Sydney," does it know this is wrong? Models assert misconceptions with the same fluency as facts, so the question ca…