4 papers
The Query Knows What to Forget: A Second Erase Direction for Linear Attention
Dhruman Gupta, Aritra Das, Debayan Gupta
Linear attention keeps a state of fixed size. At long context, many stored items share this state, and interference between them degrades retrieval. Gated DeltaNet-2 (GDN-2), like…
Linearized 2-Simplicial Attention
Aritra Das, Dhruman Gupta, Debayan Gupta
We present a linearized form of 2-simplicial attention by rewriting the trilinear score as an inner product between a composite query and a key, so that the sum over one token axis…
Generating DDPM-based Samples from Tilted Distributions
Himadri Mandal, Dhruman Gupta, Rushil Gupta +5
Given independent samples from a -dimensional probability distribution, our aim is to generate diffusion-based samples from a distribution obtained by tilting the original,…
Fundamental limits for weighted empirical approximations of tilted distributions
Sarvesh Ravichandran Iyer, Himadri Mandal, Dhruman Gupta +5
Consider the task of generating samples from a tilted distribution of a random vector whose underlying distribution is unknown, but samples from it are available. This finds applic…