3 papers
cs.LG2026
SharedSAE: One Feature Dictionary Across Language Models
Daniil Ognev, Célian Vasson, Lijie Hu +2
Sparse autoencoders (SAEs) are widely used to interpret language model activations, but SAE training and latent labelling are typically repeated for every model. Here, we show that…
cs.LG2026
Value-Gradient Hypothesis of RL for LLMs
Arip Asadulaev, Daniil Ognev, Karim Salta +1
Reinforcement learning substantially improves pretrained language models, but it remains understudied why critic-free methods such as PPO and GRPO work as well as they do, and when…
cs.CR2026
Adaptively Robust LLM Monitoring via Activation Watermarking
Toluwani Aremu, Daniil Ognev, Samuele Poppi +1
Providers monitor deployed large language models (LLMs) to detect misuse that they cannot prevent. LLM monitoring is deterministic and often openly available, so …