1 citations · 2 across the 3 of their papers we have counts for
3 papers
cs.LG2024★ 1 cited
Revisiting SMoE Language Models by Evaluating Inefficiencies with Task Specific Expert Pruning
Soumajyoti Sarkar, Leonard Lausen, Volkan Cevher +3
Sparse Mixture of Expert (SMoE) models have emerged as a scalable alternative to dense models in language modeling. These models use conditionally activated feedforward subnetworks…
cs.LG2024
EMC: Efficient MCMC Negative Sampling for Contrastive Learning with Global Convergence
Chung-Yiu Yau, Hoi-To Wai, Parameswaran Raman +2
A key challenge in contrastive learning is to generate negative samples from a large sample set to contrast with positive samples, for learning better encoding of the data. These n…
cs.CL2023★ 1 cited
Testing the Limits of Unified Sequence to Sequence LLM Pretraining on Diverse Table Data Tasks
Soumajyoti Sarkar, Leonard Lausen
Tables stored in databases and tables which are present in web pages and articles account for a large part of semi-structured data that is available on the internet. It then become…