12 papers
Emergent Hierarchical Monosemantic Neurons from the Group-Contrastive Forward-Forward Algorithm
Yiming Tang, Qinglin Qi, Zhaoqian Yao +2
Mechanistic interpretability has made significant strides in understanding neural network representations, with sparse dictionary learning (SDL) methods, most prominently sparse au…
Bridging Mechanistic Interpretability and Prompt Engineering with Gradient Ascent for Interpretable Persona Control
Harshvardhan Saini, Yiming Tang, Dianbo Liu
Controlling emergent behavioral personas (e.g., sycophancy, hallucination) in Large Language Models (LLMs) is critical for AI safety, yet remains a persistent challenge. Existing s…
When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models
Harshvardhan Saini, Samyak Jha, Yiming Tang +1
Vision-Language Models (VLMs) increasingly power high-stakes applications, from medical imaging to autonomous systems, yet they routinely hallucinate, confidently describing conten…
FML-bench: A Controlled Study of AI Research Agent Strategies from the Perspective of Search Dynamics
Qiran Zou, Hou Hei Lam, Wenhao Zhao +11
AI research agents accelerate ML research by automating hypothesis generation, experimentation, and empirical refinement. Existing agent strategies range from greedy hill-climbing…
CXR-LanIC: Language-Grounded Interpretable Classifier for Chest X-Ray Diagnosis
Yiming Tang, Wenjia Zhong, Rushi Shah +1
Deep learning models have achieved remarkable accuracy in chest X-ray diagnosis, yet their widespread clinical adoption remains limited by the black-box nature of their predictions…
A Unified Theory of Sparse Dictionary Learning in Mechanistic Interpretability: Piecewise Biconvexity and Spurious Minima
Yiming Tang, Harshvardhan Saini, Zhaoqian Yao +6
As AI models achieve remarkable capabilities across diverse domains, understanding what representations they learn and how they encode concepts has become increasingly important fo…