7 papers
Detection Without Correction: A Robust Asymmetry in Activation-Based Hallucination Probing
Dip Roy, Rajiv Misra, Sanjay Kumar Singh +1
Activation-based linear probing is widely proposed as a method for both detecting and correcting hallucinations in autoregressive language models. We present an empirical study acr…
Posterior-Calibrated Causal Circuits in Variational Autoencoders: Why Image-Domain Interpretability Fails on Tabular Data
Dip Roy, Rajiv Misra, Sanjay Kumar Singh +1
Although mechanism-based interpretability has generated an abundance of insight for discriminative network analysis, generative models are less understood -- particularly outside o…
A Multi-Level Causal Intervention Framework for Mechanistic Interpretability in Variational Autoencoders
Dip Roy, Rajiv Misra, Sanjay Kumar Singh +1
Understanding how generative models represent and transform data is a foundational problem in deep learning interpretability. While mechanistic interpretability of discriminative a…
MemGuard-Alpha: Detecting and Filtering Memorization-Contaminated Signals in LLM-Based Financial Forecasting via Membership Inference and Cross-Model Disagreement
Anisha Roy, Dip Roy
Large language models (LLMs) are increasingly used to generate financial alpha signals, yet growing evidence shows that LLMs memorize historical financial data from their training…
Fundamental Limits of Neural Network Sparsification: Evidence from Catastrophic Interpretability Collapse
Dip Roy, Rajiv Misra, Sanjay Kumar Singh
Extreme neural network sparsification (90% activation reduction) presents a critical challenge for mechanistic interpretability: understanding whether interpretable features surviv…
Mechanistic Interpretability of Diffusion Models: Circuit-Level Analysis and Causal Validation
Dip Roy
We present a quantitative circuit-level analysis of diffusion models, establishing computational pathways and mechanistic principles underlying image generation processes. Through…