3 papers
cs.LG2024
Beyond Parameter Count: Implicit Bias in Soft Mixture of Experts
Youngseog Chung, Dhruv Malik, Jeff Schneider +2
The traditional viewpoint on Sparse Mixture of Experts (MoE) models is that instead of training a single large expert, which is computationally expensive, we can train many small e…
cs.LG2024★ 1 cited
Role of Locality and Weight Sharing in Image-Based Tasks: A Sample Complexity Separation between CNNs, LCNs, and FCNs
Aakash Lahoti, Stefani Karp, Ezra Winston +2
Vision tasks are characterized by the properties of locality and translation invariance. The superior performance of convolutional neural networks (CNNs) on these tasks is widely a…
stat.ML2022
Complete Policy Regret Bounds for Tallying Bandits
Dhruv Malik, Yuanzhi Li, Aarti Singh
Policy regret is a well established notion of measuring the performance of an online learning algorithm against an adaptive adversary. We study restrictions on the adversary that e…