4 papers
LUNA: Linear Universal Neural Attention with Generalization Guarantees
Ashkan Shahbazi, Ping He, Ali Abbasi +6
Scaling attention faces a critical bottleneck: the quadratic computational cost of softmax attention, which limits its application in long-sequence domains. Whil…
Fused Partial Gromov-Wasserstein for Structured Objects
Yikun Bai, Shuang Wang, Huy Tran +3
Structured data, such as graphs, is vital in machine learning due to its capacity to capture complex relationships and interactions. In recent years, the Fused Gromov-Wasserstein (…
Understanding Learning with Sliced-Wasserstein Requires Rethinking Informative Slices
Huy Tran, Yikun Bai, Ashkan Shahbazi +2
The practical applications of Wasserstein distances (WDs) are constrained by their sample and computational complexities. Sliced-Wasserstein distances (SWDs) provide a workaround b…
Linear Spherical Sliced Optimal Transport: A Fast Metric for Comparing Spherical Data
Xinran Liu, Yikun Bai, Rocío Díaz Martín +5
Efficient comparison of spherical probability distributions becomes important in fields such as computer vision, geosciences, and medicine. Sliced optimal transport distances, such…