60 citations · 85 across the 3 of their papers we have counts for
4 papers · 1 filter
HyperRouter: Towards Efficient Training and Inference of Sparse Mixture of Experts
Giang Do, Khiem Le, Quang Pham +7
By routing input tokens to only a few split experts, Sparse Mixture-of-Experts has enabled efficient training of large language models. Recent findings suggest that fixing the rout…
Continual Normalization: Rethinking Batch Normalization for Online Continual Learning
Quang Pham, Chenghao Liu, Steven Hoi
Existing continual learning methods use Batch Normalization (BN) to facilitate training and improve generalization across tasks. However, the non-i.i.d and non-stationary nature of…
DualNet: Continual Learning, Fast and Slow
Quang Pham, Chenghao Liu, Steven Hoi
According to Complementary Learning Systems (CLS) theory~\citep{mcclelland1995there} in neuroscience, humans do effective \emph{continual learning} through two complementary system…
Bilevel Continual Learning
Quang Pham, Doyen Sahoo, Chenghao Liu +1
Continual learning aims to learn continuously from a stream of tasks and data in an online-learning fashion, being capable of exploiting what was learned previously to improve curr…