3 papers
cs.CL2026
Detection Without Correction: A Robust Asymmetry in Activation-Based Hallucination Probing
Dip Roy, Rajiv Misra, Sanjay Kumar Singh +1
Activation-based linear probing is widely proposed as a method for both detecting and correcting hallucinations in autoregressive language models. We present an empirical study acr…
cs.LG2026
Posterior-Calibrated Causal Circuits in Variational Autoencoders: Why Image-Domain Interpretability Fails on Tabular Data
Dip Roy, Rajiv Misra, Sanjay Kumar Singh +1
Although mechanism-based interpretability has generated an abundance of insight for discriminative network analysis, generative models are less understood -- particularly outside o…
cs.LG2026
A Multi-Level Causal Intervention Framework for Mechanistic Interpretability in Variational Autoencoders
Dip Roy, Rajiv Misra, Sanjay Kumar Singh +1
Understanding how generative models represent and transform data is a foundational problem in deep learning interpretability. While mechanistic interpretability of discriminative a…