From the 1 of 13 linked papers with an AI index.
9 papers · 1 filter
Outcome Rewards Do Not Guarantee Verifiable or Causally Important Reasoning
Qinan Yu, Alexa Tartaglini, Peter Hase +2
Reinforcement Learning from Verifiable Rewards (RLVR) on chain-of-thought reasoning has become a standard part of language model post-training recipes. A common assumption is that…
Mechanistic evaluation of Transformers and state space models
Aryaman Arora, Neil Rathi, Nikil Roashan Selvam +3
State space models (SSMs) for language modelling promise an efficient and performant alternative to quadratic-attention Transformers, yet show variable performance on recalling bas…
Bayesian scaling laws for in-context learning
Aryaman Arora, Dan Jurafsky, Christopher Potts +1
In-context learning (ICL) is a powerful technique for getting language models to perform complex tasks with no training updates. Prior work has established strong correlations betw…
HyperSteer: Activation Steering at Scale with Hypernetworks
Jiuding Sun, Sidharth Baskaran, Zhengxuan Wu +3
Steering language models (LMs) by modifying internal activations is a popular approach for controlling text generation. Unsupervised dictionary learning methods, e.g., sparse autoe…
HyperDAS: Towards Automating Mechanistic Interpretability with Hypernetworks
Jiuding Sun, Jing Huang, Sidharth Baskaran +4
Mechanistic interpretability has made great strides in identifying neural network features (e.g., directions in hidden activation space) that mediate concepts(e.g., the birth year…
AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders
Zhengxuan Wu, Aryaman Arora, Atticus Geiger +5
Fine-grained steering of language model outputs is essential for safety and reliability. Prompting and finetuning are widely used to achieve these goals, but interpretability resea…