6 papers
Geometric and Stochastic Analysis of Discontinuities in Sparse Mixture-of-Experts
Tho Tran Huu, Huu-Tuan Nguyen, Thien-Hai Nguyen +4
Sparse Mixture-of-Experts (SMoE) architectures are now widely deployed in state-of-the-art language and vision models, where conditional routing allows scaling to very large networ…
Equivariant Polynomial Functional Networks
Thieu N. Vo, Viet-Hoang Tran, Tho Tran Huu +5
Neural Functional Networks (NFNs) have gained increasing interest due to their wide range of applications, including extracting information from implicit representations of data, e…
Tree-Sliced Wasserstein Distance: A Geometric Perspective
Viet-Hoang Tran, Trang Pham, Tho Tran +4
Many variants of Optimal Transport (OT) have been developed to address its heavy computation. Among them, notably, Sliced Wasserstein (SW) is widely used for application domains by…
A Clifford Algebraic Approach to E(n)-Equivariant High-order Graph Neural Networks
Viet-Hoang Tran, Thieu N. Vo, Tho Tran Huu +1
Designing neural network architectures that can handle data symmetry is crucial. This is especially important for geometric graphs whose properties are equivariance under Euclidean…
Equivariant Neural Functional Networks for Transformers
Viet-Hoang Tran, Thieu N. Vo, An Nguyen The +5
This paper systematically explores neural functional networks (NFN) for transformer architectures. NFN are specialized neural networks that treat the weights, gradients, or sparsit…
Revisiting Kernel Attention with Correlated Gaussian Process Representation
Long Minh Bui, Tho Tran Huu, Duy Dinh +2
Transformers have increasingly become the de facto method to model sequential data with state-of-the-art performance. Due to its widespread use, being able to estimate and calibrat…