3 papers
cs.LG2026
Mixture-of-Control: State-Aware Fine-Tuning for Transformer-based Models
Duc Anh Nguyen, Tien Ngoc Luu, Tung Pham +1
State-based fine-tuning has emerged as a compelling alternative to weight-based adaptation for transformers, updating lightweight controls into states rather than model weights, of…
cs.CL2025
ClozeMath: Improving Mathematical Reasoning in Language Models by Learning to Fill Equations
Quang Hieu Pham, Thuy Duong Nguyen, Tung Pham +2
The capabilities of large language models (LLMs) have been enhanced by training on data that reflects human thought processes, such as the Chain-of-Thought format. However, evidenc…
cs.LG2025
Demystifying the Token Dynamics of Deep Selective State Space Models
Thieu N Vo, Tung D. Pham, Xin T. Tong +1
Selective state space models (SSM), such as Mamba, have gained prominence for their effectiveness in modeling sequential data. Despite their outstanding empirical performance, a co…