interpretability 1neural networks 1out-of-distribution detection 1representation learning 1sparse autoencoders 1
From the 1 of 2 linked papers with an AI index.
2 papers
cs.LG2026
Sparse Autoencoders for Interpretable Out-of-Distribution Detection
Ayush Karmacharya, Luke Luschwitz, Lucia Romero +2
The paper proposes using sparse autoencoders to extract interpretable sparse features from intermediate neural network layers and defines an OOD detection score based on cosine sim…
cs.AI2026
Subliminal Clocks: Latent Time Modelling in Diffusion Language Models
Maximo Rulli, Maximo Eduardo Rulli, Thomas Vaitses Fontanari +11
Diffusion Language Models (DLMs) have recently emerged as a promising alternative to autoregressive models. Unlike standard diffusion-based approaches, DLMs are not explicitly cond…