Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
SharedSAE: One Feature Dictionary Across Language Models
Daniil Ognev, Célian Vasson, Lijie Hu +2
Sparse autoencoders (SAEs) are widely used to interpret language model activations, but SAE training and latent labelling are typically repeated for every model. Here, we show that…
cs.LG2026
Value-Gradient Hypothesis of RL for LLMs
Arip Asadulaev, Daniil Ognev, Karim Salta +1
Reinforcement learning substantially improves pretrained language models, but it remains understudied why critic-free methods such as PPO and GRPO work as well as they do, and when…