12 papers
Nectar: Neural Estimation of Cached-Token Attention via Regression
João Monteiro, Michal Klein, Pierre Ablin +1
Evaluating softmax attention over a fixed long context requires reading every cached key-value pair for each new query token. For a given context (a book, a manual, a legal corpus)…
Amortized Maximum Inner Product Search with Learned Support Functions
Theo X. Olausson, João Monteiro, Michal Klein +1
Maximum inner product search (MIPS) is a crucial subroutine in machine learning, requiring the identification of a vector taken within a database (the keys) that best aligns with a…
The Coupling Within: Flow Matching via Distilled Normalizing Flows
David Berthelot, Tianrong Chen, Jiatao Gu +6
Flow models have rapidly become the go-to method for training and deploying large-scale generators, owing their success to inference-time flexibility via adjustable integration ste…
Flow Matching with Semidiscrete Couplings
Alireza Mousavi-Hosseini, Stephen Y. Zhang, Michal Klein +1
Flow models parameterized as time-dependent velocity fields can generate data from noise by integrating an ODE. These models are often trained using flow matching, i.e. by sampling…
LinEAS: End-to-end Learning of Activation Steering with a Distributional Loss
Pau Rodriguez, Michal Klein, Eleonora Gualdoni +5
The growing use of generative models in daily life calls for efficient mechanisms to control their generation, to e.g., produce safe content or provide users with tools to explore…
On Fitting Flow Models with Large Sinkhorn Couplings
Stephen Zhang, Alireza Mousavi-Hosseini, Michal Klein +1
Flow models transform data gradually from one modality (e.g. noise) onto another (e.g. images). Such models are parameterized by a time-dependent velocity field, trained to fit seg…