3 papers
cs.LG2026
Mixture-of-Control: State-Aware Fine-Tuning for Transformer-based Models
Duc Anh Nguyen, Tien Ngoc Luu, Tung Pham +1
State-based fine-tuning has emerged as a compelling alternative to weight-based adaptation for transformers, updating lightweight controls into states rather than model weights, of…
cs.LG2026
Selective Sinkhorn Routing for Improved Sparse Mixture of Experts
Duc Anh Nguyen, Huu Binh Ta, Nhuan Le Duc +2
Sparse Mixture-of-Experts (SMoE) models are scalable and computationally efficient, enabling large increases in model capacity with limited inference overhead. Existing SMoE method…
cs.LG2025
Generalization Bounds for Robust Contrastive Learning: From Theory to Practice
Ngoc N. Tran, Lam Tran, Hoang Phan +5
Contrastive Learning first extracts features from unlabeled data, followed by linear probing with labeled data. Adversarial Contrastive Learning (ACL) integrates Adversarial Traini…