3 papers
cs.LG2026
Soft Contamination Means Benchmarks Test Shallow Generalization
Ari Spiesberger, Juan J. Vazquez, Nicky Pochinkov +4
If LLM training data is polluted with benchmark test data, then benchmark performance gives biased estimates of out-of-distribution (OOD) generalization. Typical decontamination fi…
cs.CL2025
Beyond Tokens in Language Models: Interpreting Activations through Text Genre Chunks
Ãloïse Benito-Rodriguez, Einar Urdshals, Jasmina Nasufi +1
Understanding Large Language Models (LLMs) is key to ensure their safe and beneficial deployment. This task is complicated by the difficulty of interpretability of LLM structures,…
cs.CL2025
ParaScopes: What do Language Models Activations Encode About Future Text?
Nicky Pochinkov, Yulia Volkova, Anna Vasileva +1
Interpretability studies in language models often investigate forward-looking representations of activations. However, as language models become capable of doing ever longer time h…