3 papers
cs.LG2026
Building Fast, Evaluating Slow: Pipeline Choices Dominate Autointerpretability Score Variance
Sinie van der Ben, Neele Roch, Anna Hedström +1
Cross-paper comparison of sparse autoencoder (SAE) interpretability often relies on autointerpretability scores. In this evaluation pipeline, a language model (LM) explains each fe…
cs.CL2026
Where Do Models Find Happiness? Emotion Vectors in Open-Source LLMs
Sinie van der Ben, Raphaël Baur, Yannick Metz +1
Recent work identified emotion vectors in Claude Sonnet 4.5, which are internal representations that encode emotion concepts, causally influence behavior, and exhibit geometry mirr…
cs.CG2025
Explainable Mapper: Charting LLM Embedding Spaces Using Perturbation-Based Explanation and Verification Agents
Xinyuan Yan, Rita Sevastjanova, Sinie van der Ben +2
Large language models (LLMs) produce high-dimensional embeddings that capture rich semantic and syntactic relationships between words, sentences, and concepts. Investigating the to…