7 papers
LUNA: Linear Universal Neural Attention with Generalization Guarantees
Ashkan Shahbazi, Ping He, Ali Abbasi +6
Scaling attention faces a critical bottleneck: the quadratic computational cost of softmax attention, which limits its application in long-sequence domains. Whil…
EMPEROR: Efficient Moment-Preserving Representation of Distributions
Xinran Liu, Shansita D. Sharma, Soheil Kolouri
We introduce EMPEROR (Efficient Moment-Preserving Representation of Distributions), a mathematically rigorous and computationally efficient framework for representing high-dimensio…
Constrained Sliced Wasserstein Embedding
Navid NaderiAlizadeh, Darian Salehi, Xinran Liu +1
Sliced Wasserstein (SW) distances offer an efficient method for comparing high-dimensional probability measures by projecting them onto multiple 1-dimensional probability distribut…
Policy Search, Retrieval, and Composition via Task Similarity in Collaborative Agentic Systems
Saptarshi Nath, Christos Peridis, Eseoghene Benjamin +7
Agentic AI aims to create systems that set their own goals, adapt proactively to change, and refine behavior through continuous experience. Recent advances suggest that, when facin…
Fused Partial Gromov-Wasserstein for Structured Objects
Yikun Bai, Shuang Wang, Huy Tran +3
Structured data, such as graphs, is vital in machine learning due to its capacity to capture complex relationships and interactions. In recent years, the Fused Gromov-Wasserstein (…
ESPFormer: Doubly-Stochastic Attention with Expected Sliced Transport Plans
Ashkan Shahbazi, Elaheh Akbari, Darian Salehi +3
While self-attention has been instrumental in the success of Transformers, it can lead to over-concentration on a few tokens during training, resulting in suboptimal information fl…