collaborators

6 papers

cs.LG2026

Geometric and Stochastic Analysis of Discontinuities in Sparse Mixture-of-Experts

Tho Tran Huu, Huu-Tuan Nguyen, Thien-Hai Nguyen +4

Sparse Mixture-of-Experts (SMoE) architectures are now widely deployed in state-of-the-art language and vision models, where conditional routing allows scaling to very large networ…

cs.LG2025

Equivariant Polynomial Functional Networks

Thieu N. Vo, Viet-Hoang Tran, Tho Tran Huu +5

Neural Functional Networks (NFNs) have gained increasing interest due to their wide range of applications, including extracting information from implicit representations of data, e…

cs.LG2025

Tree-Sliced Wasserstein Distance: A Geometric Perspective

Viet-Hoang Tran, Trang Pham, Tho Tran +4

Many variants of Optimal Transport (OT) have been developed to address its heavy computation. Among them, notably, Sliced Wasserstein (SW) is widely used for application domains by…

cs.LG2025

A Clifford Algebraic Approach to E(n)-Equivariant High-order Graph Neural Networks

Viet-Hoang Tran, Thieu N. Vo, Tho Tran Huu +1

Designing neural network architectures that can handle data symmetry is crucial. This is especially important for geometric graphs whose properties are equivariance under Euclidean…

cs.LG2025

Equivariant Neural Functional Networks for Transformers

Viet-Hoang Tran, Thieu N. Vo, An Nguyen The +5

This paper systematically explores neural functional networks (NFN) for transformer architectures. NFN are specialized neural networks that treat the weights, gradients, or sparsit…

cs.LG2025

Revisiting Kernel Attention with Correlated Gaussian Process Representation

Long Minh Bui, Tho Tran Huu, Duy Dinh +2

Transformers have increasingly become the de facto method to model sequential data with state-of-the-art performance. Due to its widespread use, being able to estimate and calibrat…