6 papers
Detection Without Correction: A Robust Asymmetry in Activation-Based Hallucination Probing
Dip Roy, Rajiv Misra, Sanjay Kumar Singh +1
Activation-based linear probing is widely proposed as a method for both detecting and correcting hallucinations in autoregressive language models. We present an empirical study acr…
Posterior-Calibrated Causal Circuits in Variational Autoencoders: Why Image-Domain Interpretability Fails on Tabular Data
Dip Roy, Rajiv Misra, Sanjay Kumar Singh +1
Although mechanism-based interpretability has generated an abundance of insight for discriminative network analysis, generative models are less understood -- particularly outside o…
A Multi-Level Causal Intervention Framework for Mechanistic Interpretability in Variational Autoencoders
Dip Roy, Rajiv Misra, Sanjay Kumar Singh +1
Understanding how generative models represent and transform data is a foundational problem in deep learning interpretability. While mechanistic interpretability of discriminative a…
Fundamental Limits of Neural Network Sparsification: Evidence from Catastrophic Interpretability Collapse
Dip Roy, Rajiv Misra, Sanjay Kumar Singh
Extreme neural network sparsification (90% activation reduction) presents a critical challenge for mechanistic interpretability: understanding whether interpretable features surviv…
HinTel-AlignBench: A Framework and Benchmark for Hindi-Telugu with English-Aligned Samples
Rishikant Chigrupaatii, Ponnada Sai Tulasi Kanishka, Lalit Chandra Routhu +6
With nearly 1.5 billion people and more than 120 major languages, India represents one of the most diverse regions in the world. As multilingual Vision-Language Models (VLMs) gain…
GASCADE: Grouped Summarization of Adverse Drug Event for Enhanced Cancer Pharmacovigilance
Sofia Jamil, Aryan Dabad, Bollampalli Areen Reddy +3
In the realm of cancer treatment, summarizing adverse drug events (ADEs) reported by patients using prescribed drugs is crucial for enhancing pharmacovigilance practices and improv…