4 papers
Emergent Hierarchical Monosemantic Neurons from the Group-Contrastive Forward-Forward Algorithm
Yiming Tang, Qinglin Qi, Zhaoqian Yao +2
Mechanistic interpretability has made significant strides in understanding neural network representations, with sparse dictionary learning (SDL) methods, most prominently sparse au…
Bridging Mechanistic Interpretability and Prompt Engineering with Gradient Ascent for Interpretable Persona Control
Harshvardhan Saini, Yiming Tang, Dianbo Liu
Controlling emergent behavioral personas (e.g., sycophancy, hallucination) in Large Language Models (LLMs) is critical for AI safety, yet remains a persistent challenge. Existing s…
When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models
Harshvardhan Saini, Samyak Jha, Yiming Tang +1
Vision-Language Models (VLMs) increasingly power high-stakes applications, from medical imaging to autonomous systems, yet they routinely hallucinate, confidently describing conten…
A Unified Theory of Sparse Dictionary Learning in Mechanistic Interpretability: Piecewise Biconvexity and Spurious Minima
Yiming Tang, Harshvardhan Saini, Zhaoqian Yao +6
As AI models achieve remarkable capabilities across diverse domains, understanding what representations they learn and how they encode concepts has become increasingly important fo…