11 citations · 19 across the 2 of their papers we have counts for
2 papers
cs.LG2022★ 8 cited
Sparsely Activated Mixture-of-Experts are Robust Multi-Task Learners
Shashank Gupta, Subhabrata Mukherjee, Krishan Subudhi +4
Traditional multi-task learning (MTL) methods use dense networks that use the same set of shared weights across several different tasks. This often creates interference where two o…
cs.CL2022★ 11 cited
Knowledge Infused Decoding
Ruibo Liu, Guoqing Zheng, Shashank Gupta +5
Pre-trained language models (LMs) have been shown to memorize a substantial amount of knowledge from the pre-training corpora; however, they are still limited in recalling factuall…