7 papers
The Weight Gram Matrix Captures Sequential Feature Linearization in Deep Networks
Taehun Cha, Daniel Beaglehole, Adityanarayanan Radhakrishnan +1
Understanding how deep neural networks learn representations remains a central challenge in machine learning theory. In this work, we propose a feature-centric framework for analyz…
Contextual Linear Activation Steering of Language Models
Brandon Hsu, Daniel Beaglehole, Adityanarayanan Radhakrishnan +1
Linear activation steering is a powerful approach for eliciting the capabilities of large language models and specializing their behavior using limited labeled data. While effectiv…
xRFM: Accurate, scalable, and interpretable feature learning models for tabular data
Daniel Beaglehole, David Holzmüller, Adityanarayanan Radhakrishnan +1
Inference from tabular data, collections of continuous and categorical variables organized into matrices, is a foundation for modern technology and science. Yet, in contrast to the…
Efficient and accurate steering of Large Language Models through attention-guided feature learning
Parmida Davarmanesh, Ashia Wilson, Adityanarayanan Radhakrishnan
Steering, or direct manipulation of internal activations to guide LLM responses toward specific semantic concepts, is emerging as a promising avenue for both understanding how sema…
Emergence in non-neural models: grokking modular arithmetic via average gradient outer product
Neil Mallinar, Daniel Beaglehole, Libin Zhu +3
Neural networks trained to solve modular arithmetic tasks exhibit grokking, a phenomenon where the test accuracy starts improving long after the model achieves 100% training accura…
Toward universal steering and monitoring of AI models
Daniel Beaglehole, Adityanarayanan Radhakrishnan, Enric Boix-Adserà +1
Modern AI models contain much of human knowledge, yet understanding of their internal representation of this knowledge remains elusive. Characterizing the structure and properties…