12 papers
Emergent Hierarchical Monosemantic Neurons from the Group-Contrastive Forward-Forward Algorithm
Yiming Tang, Qinglin Qi, Zhaoqian Yao +2
Mechanistic interpretability has made significant strides in understanding neural network representations, with sparse dictionary learning (SDL) methods, most prominently sparse au…
Bridging Mechanistic Interpretability and Prompt Engineering with Gradient Ascent for Interpretable Persona Control
Harshvardhan Saini, Yiming Tang, Dianbo Liu
Controlling emergent behavioral personas (e.g., sycophancy, hallucination) in Large Language Models (LLMs) is critical for AI safety, yet remains a persistent challenge. Existing s…
When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models
Harshvardhan Saini, Samyak Jha, Yiming Tang +1
Vision-Language Models (VLMs) increasingly power high-stakes applications, from medical imaging to autonomous systems, yet they routinely hallucinate, confidently describing conten…
JEDI: Joint Embedding Diffusion World Model for Online Model-Based Reinforcement Learning
Jing Yu Lim, Rushi Shah, Zarif Ikram +4
Diffusion world models have recently become competitive for online model-based reinforcement learning, but current approaches expose a tension: pixel diffusion is effective but com…
AB-Sparse: Sparse Attention with Adaptive Block Size for Accurate and Efficient Long-Context Inference
Di Liu, Ruitian Wang, Chen Chen +6
As large language models scale to longer contexts, loading the growing KV cache during attention computation becomes a critical bottleneck. Previous work has shown that attention c…
Absurd World: A Simple Yet Powerful Method to Absurdify the Real-world for Probing LLM Reasoning Capabilities
Ryan Albright, Golam Md Muktadir, Zarif Ikram +3
While extremely powerful and versatile at various tasks, the thinking capabilities of large language models (LLMs) are often put under scrutiny as they sometimes fail to solve prob…