35 papers
The Noetherian Case of Bayart's Power-Series Question
Viet-Hoang Tran, Dung V. Nguyen, Quang X. Nguyen +2
Let be a commutative Noetherian ring. We prove that if the one-variable formal power-series ring is a unique factorization domain, then so is the two-variable formal p…
MuonSSM: Orthogonalizing State Space Models for Sequence Modeling
Thai-Khanh Nguyen, Ngoc-Bich-Uyen Vo, Thieu N. Vo +2
State space models (SSMs) have emerged as efficient linear-time alternatives to attention for long-sequence modeling. However, existing SSMs often suffer from instability and memor…
Geometric and Stochastic Analysis of Discontinuities in Sparse Mixture-of-Experts
Tho Tran Huu, Huu-Tuan Nguyen, Thien-Hai Nguyen +4
Sparse Mixture-of-Experts (SMoE) architectures are now widely deployed in state-of-the-art language and vision models, where conditional routing allows scaling to very large networ…
Functional Equivalence in Attention: A Comprehensive Study with Applications to Linear Mode Connectivity
Viet-Hoang Tran, Vinh Khanh Bui, Van-Hoan Trinh +2
Neural network parameter spaces are inherently non-injective, as distinct parameter configurations can realize identical functions through functional equivalence. While this symmet…
Conservation Laws for Modern Neural Architectures
Viet-Hoang Tran, Vinh Khanh Bui, Tan Lai Ngoc +3
Understanding gradient descent dynamics is key to explaining the success of over-parameterized models, where implicit bias manifests through conservation laws in gradient flow. Whi…
High-Dimensional Random Projection for Activation Steering in Language Models
Minh-Hieu Pham, Bach Do, Laziz Abdullaev +2
Activation steering has emerged as a key methodology for controlling the behavior of large language models (LLMs). Existing difference-in-means based methods, however, are fundamen…