6 papers
Dynamic Expert Sharing: Decoupling Memory from Parallelism in Mixture-of-Experts Diffusion LLMs
Hao Mark Chen, Zhiwen Mo, Royson Lee +6
Among parallel decoding paradigms, diffusion large language models (dLLMs) have emerged as a promising candidate that balances generation quality and throughput. However, their int…
The Reward Model Selection Crisis in Personalized Alignment
Fady Rezk, Yuangang Pan, Chuan-Sheng Foo +4
Personalized alignment from preference data has focused primarily on improving personal reward model (RM) accuracy, with the implicit assumption that better preference ranking tran…
FW-Merging: Scaling Model Merging with Frank-Wolfe Optimization
Hao Mark Chen, Shell Xu Hu, Wayne Luk +2
Model merging has emerged as a promising approach for multi-task learning (MTL), offering a data-efficient alternative to conventional fine-tuning. However, with the rapid developm…
Model Diffusion for Certifiable Few-shot Transfer Learning
Fady Rezk, Royson Lee, Henry Gouk +2
In contemporary deep learning, a prevalent and effective workflow for solving low-data problems is adapting powerful pre-trained foundation models (FMs) to new tasks via parameter-…
FedPEFT: Federated Learning to Personalize PEFT for Multilingual LLMs
Royson Lee, Minyoung Kim, Fady Rezk +3
Federated learning (FL) has enabled the training of multilingual large language models (LLMs) on diverse and decentralized multilingual data, especially on low-resource languages.…
A Bayesian Approach to Data Point Selection
Xinnuo Xu, Minyoung Kim, Royson Lee +2
Data point selection (DPS) is becoming a critical topic in deep learning due to the ease of acquiring uncurated training data compared to the difficulty of obtaining curated or pro…