4 papers
Learning Correlated Reward Models: Statistical Barriers and Opportunities
Yeshwanth Cherapanamjeri, Constantinos Daskalakis, Gabriele Farina +1
Random Utility Models (RUMs) are a classical framework for modeling user preferences and play a key role in reward modeling for Reinforcement Learning from Human Feedback (RLHF). H…
The Space Complexity of Learning-Unlearning Algorithms
Yeshwanth Cherapanamjeri, Sumegha Garg, Nived Rajaraman +2
We study the memory complexity of machine unlearning algorithms that provide strong data deletion guarantees to the users. Formally, consider an algorithm for a particular learning…
Efficient Automated Circuit Discovery in Transformers using Contextual Decomposition
Aliyah R. Hsu, Georgia Zhou, Yeshwanth Cherapanamjeri +4
Automated mechanistic interpretation research has attracted great interest due to its potential to scale explanations of neural network internals to large models. Existing automate…
Heavy-tailed Contamination is Easier than Adversarial Contamination
Yeshwanth Cherapanamjeri, Daniel Lee
A large body of work in the statistics and computer science communities dating back to Huber (Huber, 1960) has led to statistically and computationally efficient outlier-robust est…