2 papers
cs.LG2024
Learning to Route Among Specialized Experts for Zero-Shot Generalization
Mohammed Muqeeth, Haokun Liu, Yufan Liu +1
Recently, there has been a widespread proliferation of "expert" language models that are specialized to a specific task or domain through parameter-efficient fine-tuning. How can w…
cs.LG2024
Soft Merging of Experts with Adaptive Routing
Mohammed Muqeeth, Haokun Liu, Colin Raffel
Sparsely activated neural networks with conditional computation learn to route their inputs through different "expert" subnetworks, providing a form of modularity that densely acti…