39 citations · 82 across the 19 of their papers we have counts for
Showing 2025 · cs.CLShow all
2 papers · 2 filters
cs.CL2025
Calibrating Lightweight Sparse Autoencoder Feature Steering
Ananya Joshi, Celia Cintas, Skyler Speakman
Sparse autoencoders (SAEs) can enable inference-time topic steering by modifying latent feature activations, but existing steering methods often fail when target-aligned features a…
cs.CL2025
Localizing Persona Representations in LLMs
Celia Cintas, Miriam Rateike, Erik Miehling +2
We present a study on how and where personas -- defined by distinct sets of human characteristics, values, and beliefs -- are encoded in the representation space of large language…