3 papers
cs.LG2025
Self-Adapting Language Models
Adam Zweiger, Jyothish Pari, Han Guo +3
Large language models (LLMs) are powerful but static; they lack mechanisms to adapt their weights in response to new tasks, knowledge, or examples. We introduce Self-Adapting LLMs…
cs.LG2025
On the Duality between Gradient Transformations and Adapters
Lucas Torroba-Hennigen, Hunter Lang, Han Guo +1
We study memory-efficient optimization of neural networks (in particular language models) with linear gradient transformations, where the gradients are linearly mapped to a lower d…
cs.AI2024
The Surprising Effectiveness of Test-Time Training for Few-Shot Learning
Ekin Akyürek, Mehul Damani, Adam Zweiger +5
Language models (LMs) have shown impressive performance on tasks within their training distribution, but often struggle with structurally novel tasks even when given a small number…