1 citations · 1 across the 4 of their papers we have counts for
7 papers
ReasonCACHE: Teaching LLMs To Reason Without Weight Updates
Sharut Gupta, Phillip Isola, Stefanie Jegelka +4
Can Large language models (LLMs) learn to reason without any weight update and only through in-context learning (ICL)? ICL is strikingly sample-efficient, often learning from only…
Why Less is More (Sometimes): A Theory of Data Curation
Elvis Dohmatob, Mohammad Pezeshki, Reyhane Askari-Hemmat
This paper introduces a theoretical framework to resolve a central paradox in modern machine learning: When is it better to use less data? This question has become critical as clas…
Iterative Amortized Inference: Unifying In-Context Learning and Learned Optimizers
Sarthak Mittal, Divyat Mahajan, Guillaume Lajoie +1
Modern learning systems increasingly rely on amortized learning - the idea of reusing computation or inductive biases shared across tasks to enable rapid generalization to novel pr…
Unveiling Simplicities of Attention: Adaptive Long-Context Head Identification
Konstantin Donhauser, Charles Arnal, Mohammad Pezeshki +3
The ability to process long contexts is crucial for many natural language processing tasks, yet it remains a significant challenge. While substantial progress has been made in enha…
Steering Large Language Model Activations in Sparse Spaces
Reza Bayat, Ali Rahimi-Kalahroudi, Mohammad Pezeshki +2
A key challenge in AI alignment is guiding large language models (LLMs) to follow desired behaviors at test time. Activation steering, which modifies internal model activations dur…
Improving the Scaling Laws of Synthetic Data with Deliberate Practice
Reyhane Askari-Hemmat, Mohammad Pezeshki, Elvis Dohmatob +6
Inspired by the principle of deliberate practice in human learning, we propose Deliberate Practice for Synthetic Data Generation (DP), a novel framework that improves sample effici…