3 papers
cs.LG2026
Minimizing Collateral Damage in Activation Steering
Tam Nguyen, Tu Anh Nguyen, Sina Alemohammad +1
Activation steering is a method for controlling Large Language Model (LLM) behavior by intervening in its internal representations to increase the alignment with a specific target…
cs.LG2025
Improving Routing in Sparse Mixture of Experts with Graph of Tokens
Tam Nguyen, Ngoc N. Tran, Khai Nguyen +1
Sparse Mixture of Experts (SMoE) has emerged as a key to achieving unprecedented scalability in deep learning. By activating only a small subset of parameters per sample, SMoE achi…
cs.LG2024
A Primal-Dual Framework for Transformers and Neural Networks
Tan M. Nguyen, Tam Nguyen, Nhat Ho +3
Self-attention is key to the remarkable success of transformers in sequence modeling tasks including many applications in natural language processing and computer vision. Like neur…