3 papers
cs.LG2025
Provable optimal transport with transformers: The essence of depth and prompt engineering
Hadi Daneshmand
Despite their empirical success, the internal mechanism by which transformer models align tokens during language processing remains poorly understood. This paper provides a mechani…
cs.LG2025
Linear Transformers Implicitly Discover Unified Numerical Algorithms
Patrick Lutz, Aditya Gangrade, Hadi Daneshmand +1
We train a linear attention transformer on millions of masked-block matrix completion tasks: each prompt is masked low-rank matrix whose missing block may be (i) a scalar predictio…
cs.LG2025
Data Generation without Function Estimation
Hadi Daneshmand, Ashkan Soleymani
Estimating the score function (or other population-density-dependent functions) is a fundamental component of most generative models. However, such function estimation is computati…