collaborators
Showing cs.CLShow all

6 papers · 1 filter

cs.CL2026

Outcome Rewards Do Not Guarantee Verifiable or Causally Important Reasoning

Qinan Yu, Alexa Tartaglini, Peter Hase +2

Reinforcement Learning from Verifiable Rewards (RLVR) on chain-of-thought reasoning has become a standard part of language model post-training recipes. A common assumption is that…

cs.CL2025

HyperSteer: Activation Steering at Scale with Hypernetworks

Jiuding Sun, Sidharth Baskaran, Zhengxuan Wu +3

Steering language models (LMs) by modifying internal activations is a popular approach for controlling text generation. Unsupervised dictionary learning methods, e.g., sparse autoe…

cs.CL2025

Mechanistic evaluation of Transformers and state space models

Aryaman Arora, Neil Rathi, Nikil Roashan Selvam +3

State space models (SSMs) for language modelling promise an efficient and performant alternative to quadratic-attention Transformers, yet show variable performance on recalling bas…

cs.CL2025

HyperDAS: Towards Automating Mechanistic Interpretability with Hypernetworks

Jiuding Sun, Jing Huang, Sidharth Baskaran +4

Mechanistic interpretability has made great strides in identifying neural network features (e.g., directions in hidden activation space) that mediate concepts(e.g., the birth year…

cs.CL2025

AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Zhengxuan Wu, Aryaman Arora, Atticus Geiger +5

Fine-grained steering of language model outputs is essential for safety and reliability. Prompting and finetuning are widely used to achieve these goals, but interpretability resea…

cs.CL2024

Bayesian scaling laws for in-context learning

Aryaman Arora, Dan Jurafsky, Christopher Potts +1

In-context learning (ICL) is a powerful technique for getting language models to perform complex tasks with no training updates. Prior work has established strong correlations betw…