4 papers
Diffusion-Inspired Reconfiguration of Transformers for Uncertainty Calibration
Manh Cuong Dao, Quang Hung Pham, Phi Le Nguyen +3
Uncertainty calibration in pre-trained transformers is critical for their reliable deployment in risk-sensitive applications. Yet, most existing pre-trained transformers do not hav…
Multi-Scale Finetuning for Encoder-based Time Series Foundation Models
Zhongzheng Qiao, Chenghao Liu, Yiming Zhang +6
Time series foundation models (TSFMs) demonstrate impressive zero-shot performance for time series forecasting. However, an important yet underexplored challenge is how to effectiv…
CompeteSMoE -- Statistically Guaranteed Mixture of Experts Training via Competition
Nam V. Nguyen, Huy Nguyen, Quang Pham +3
Sparse mixture of experts (SMoE) offers an appealing solution to scale up the model complexity beyond the mean of increasing the network's depth or width. However, we argue that ef…
Sequence Transferability and Task Order Selection in Continual Learning
Thinh Nguyen, Cuong N. Nguyen, Quang Pham +4
In continual learning, understanding the properties of task sequences and their relationships to model performance is important for developing advanced algorithms with better accur…