31 citations · 68 across the 22 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Expert Upcycling: Shifting the Compute-Efficient Frontier of Mixture-of-Experts
Chaitanya Dwivedi, Binxuan Huang, Himanshu Gupta +3
Mixture-of-Experts (MoE) has become the dominant architecture for scaling large language models: frontier models routinely decouple total parameters from per-token computation thro…
cs.LG2022
Let the Model Decide its Curriculum for Multitask Learning
Neeraj Varshney, Swaroop Mishra, Chitta Baral
Curriculum learning strategies in prior multi-task learning approaches arrange datasets in a difficulty hierarchy either based on human perception or by exhaustively searching the…