1 citations · 1 across the 3 of their papers we have counts for
4 papers · 1 filter
Transformers as In-Context Samplers: From Closed-Form Diffusion to Estimation-Free Sampling
Arman Adibi, Alireza Jafari, Mohammad Ghavamzadeh +1
A growing body of work establishes that large language models are not mere statistical memorizers, but are capable of in-context learning: performing inference at test time using o…
Linear Transformers Implicitly Discover Unified Numerical Algorithms
Patrick Lutz, Aditya Gangrade, Hadi Daneshmand +1
We train a linear attention transformer on millions of masked-block matrix completion tasks: each prompt is masked low-rank matrix whose missing block may be (i) a scalar predictio…
Data Generation without Function Estimation
Hadi Daneshmand, Ashkan Soleymani
Estimating the score function (or other population-density-dependent functions) is a fundamental component of most generative models. However, such function estimation is computati…
Provable optimal transport with transformers: The essence of depth and prompt engineering
Hadi Daneshmand
Despite their empirical success, the internal mechanism by which transformer models align tokens during language processing remains poorly understood. This paper provides a mechani…