3 papers
cs.LG2026
Estimating Tail Risks in Language Model Output Distributions
Rico Angell, Raghav Singhal, Zachary Horvitz +4
Language models are increasingly capable and are being rapidly deployed on a population-level scale. As a result, the safety of these models is increasingly high-stakes. Fortunatel…
cs.LG2024
Measuring Progress in Dictionary Learning for Language Model Interpretability with Board Game Models
Adam Karvonen, Benjamin Wright, Can Rager +6
What latent features are encoded in language model (LM) representations? Recent work on training sparse autoencoders (SAEs) to disentangle interpretable features in LM representati…
cs.LG2024
Polynomial Precision Dependence Solutions to Alignment Research Center Matrix Completion Problems
Rico Angell
We present solutions to the matrix completion problems proposed by the Alignment Research Center that have a polynomial dependence on the precision . The motivation fo…