collaborators

7 papers

cs.LG2026

Fast KV Compaction via Attention Matching

Adam Zweiger, Xinghong Fu, Han Guo +1

Scaling language models to long contexts is often bottlenecked by the size of the key-value (KV) cache. In deployed settings, long contexts are typically managed through compaction…

cs.LG2026

Log-Linear Attention

Han Guo, Songlin Yang, Tarushii Goel +3

The attention mechanism in Transformers is an important primitive for accurate and scalable sequence modeling. Its quadratic-compute and linear-memory complexity however remain sig…

cs.LG2025

Self-Adapting Language Models

Adam Zweiger, Jyothish Pari, Han Guo +3

Large language models (LLMs) are powerful but static; they lack mechanisms to adapt their weights in response to new tasks, knowledge, or examples. We introduce Self-Adapting LLMs…

cs.LG2025

On the Duality between Gradient Transformations and Adapters

Lucas Torroba-Hennigen, Hunter Lang, Han Guo +1

We study memory-efficient optimization of neural networks (in particular language models) with linear gradient transformations, where the gradients are linearly mapped to a lower d…

cs.AI2025

The Surprising Effectiveness of Test-Time Training for Few-Shot Learning

Ekin Akyürek, Mehul Damani, Adam Zweiger +5

Language models (LMs) have shown impressive performance on tasks within their training distribution, but often struggle with structurally novel tasks even when given a small number…

cs.CL2025

Training-Free Activation Sparsity in Large Language Models

James Liu, Pragaash Ponnusamy, Tianle Cai +3

Activation sparsity can enable practical inference speedups in large language models (LLMs) by reducing the compute and memory-movement required for matrix multiplications during t…