2 papers
cs.LG2026
Beyond a Bag of Features: Set-Level Instability in Sparse Autoencoders
Nikolai Bolik, Lennart Stöpler, Artur Andrzejak
Shani et al. (2026) show that LLM representations broadly recover human category boundaries, while failing to reflect fine-grained typicality structure. Their analysis uses cosine…
cs.CL2025
Towards Developmentally Plausible Rewards: Communicative Success as a Learning Signal for Interactive Language Models
Lennart Stöpler, Rufat Asadli, Mitja Nikolaus +2
We propose a method for training language models in an interactive setting inspired by child language acquisition. In our setting, a speaker attempts to communicate some informatio…