1 citations · 1 across the 2 of their papers we have counts for
4 papers
Flexible and Effective Mixing of Large Language Models into a Mixture of Domain Experts
Rhui Dih Lee, Laura Wynter, Raghu Kiran Ganti
We present a toolkit for creating low-cost Mixture-of-Domain-Experts (MOE) from trained models. The toolkit can be used for creating a mixture from models or from adapters. We perf…
Enhancing Training Efficiency Using Packing with Flash Attention
Achintya Kundu, Rhui Dih Lee, Laura Wynter +2
Padding is often used in tuning LLM models by adding special tokens to shorter training examples to match the length of the longest sequence in each batch. While this ensures unifo…
Rethinking Data-driven Networking with Foundation Models: Challenges and Opportunities
Franck Le, Mudhakar Srivatsa, Raghu Ganti +1
Foundational models have caused a paradigm shift in the way artificial intelligence (AI) systems are built. They have had a major impact in natural language processing (NLP), and s…
neuralRank: Searching and ranking ANN-based model repositories
Nirmit Desai, Linsong Chu, Raghu K. Ganti +2
Widespread applications of deep learning have led to a plethora of pre-trained neural network models for common tasks. Such models are often adapted from other models via transfer…