works on

From the 1 of 5 linked papers with an AI index.

collaborators

5 papers

cs.LG2026

PRISM Edit: One Vector for All Temporal Answers

Chen Huang, Qi Zheng, Ruiqin Zheng +2

The paper proposes PRISM Edit, a method that updates large language models to handle changing temporal facts by learning a single representation that can be modulated for different…

cs.CL2025

HyperSteer: Activation Steering at Scale with Hypernetworks

Jiuding Sun, Sidharth Baskaran, Zhengxuan Wu +3

Steering language models (LMs) by modifying internal activations is a popular approach for controlling text generation. Unsupervised dictionary learning methods, e.g., sparse autoe…

cs.CL2025

Improved Representation Steering for Language Models

Zhengxuan Wu, Qinan Yu, Aryaman Arora +2

Steering methods for language models (LMs) seek to provide fine-grained and interpretable control over model generations by variously changing model inputs, weights, or representat…

cs.AI2025

Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability

Atticus Geiger, Duligur Ibeling, Amir Zur +8

Causal abstraction provides a theoretical foundation for mechanistic interpretability, the field concerned with providing intelligible algorithms that are faithful simplifications…

cs.CL2025

AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Zhengxuan Wu, Aryaman Arora, Atticus Geiger +5

Fine-grained steering of language model outputs is essential for safety and reliability. Prompting and finetuning are widely used to achieve these goals, but interpretability resea…