3 papers
cs.LG2026
ASAP: Amortized Doubly-Stochastic Attention via Sliced Dual Projection
Huy Tran, Max Milkert, David Hyde
Doubly-stochastic attention has emerged as a transport-based alternative to row-softmax attention, with recent Transformer variants using it to reduce attention sinks and rank coll…
cs.LG2025
Fused Partial Gromov-Wasserstein for Structured Objects
Yikun Bai, Shuang Wang, Huy Tran +3
Structured data, such as graphs, is vital in machine learning due to its capacity to capture complex relationships and interactions. In recent years, the Fused Gromov-Wasserstein (…
cs.LG2024
Understanding Learning with Sliced-Wasserstein Requires Rethinking Informative Slices
Huy Tran, Yikun Bai, Ashkan Shahbazi +2
The practical applications of Wasserstein distances (WDs) are constrained by their sample and computational complexities. Sliced-Wasserstein distances (SWDs) provide a workaround b…