4 papers
Customizing the Inductive Biases of Softmax Attention using Structured Matrices
Yilun Kuang, Noah Amsel, Sanae Lotfi +3
The core component of attention is the scoring function, which transforms the inputs into low-dimensional queries and keys and takes the dot product of each pair. While the low-dim…
Bayesian Optimization of Antibodies Informed by a Generative Model of Evolving Sequences
Alan Nawzad Amin, Nate Gruver, Yilun Kuang +6
To build effective therapeutics, biologists iteratively mutate antibody sequences to improve binding and stability. Proposed mutations can be informed by previous measurements or b…
Unlocking Tokens as Data Points for Generalization Bounds on Larger Language Models
Sanae Lotfi, Yilun Kuang, Brandon Amos +3
Large language models (LLMs) with billions of parameters excel at predicting the next token in a sequence. Recent work computes non-vacuous compression-based generalization bounds…
Non-Vacuous Generalization Bounds for Large Language Models
Sanae Lotfi, Marc Finzi, Yilun Kuang +3
Modern language models can contain billions of parameters, raising the question of whether they can generalize beyond the training data or simply parrot their training corpora. We…