2 papers
cs.CL2026
Detection Without Correction: A Robust Asymmetry in Activation-Based Hallucination Probing
Dip Roy, Rajiv Misra, Sanjay Kumar Singh +1
Activation-based linear probing is widely proposed as a method for both detecting and correcting hallucinations in autoregressive language models. We present an empirical study acr…
cs.LG2026
A Multi-Level Causal Intervention Framework for Mechanistic Interpretability in Variational Autoencoders
Dip Roy, Rajiv Misra, Sanjay Kumar Singh +1
Understanding how generative models represent and transform data is a foundational problem in deep learning interpretability. While mechanistic interpretability of discriminative a…