1 citations · 1 across the 8 of their papers we have counts for
9 papers
To 2:4 Sparsity and Beyond: Neuron-level Activation Function to Accelerate LLM Pre-Training
Meghana Madhyastha, Daniel Haziza, Jesse Cai +3
Trainings of Large Language Models are generally bottlenecked by matrix multiplications. In the Transformer architecture, a large portion of these operations happens in the Feed Fo…
On Harnessing Idle Compute at the Edge for Foundation Model Training
Leyang Xue, Meghana Madhyastha, Myungjin Lee +3
The foundation-model ecosystem remains highly centralized because training requires immense compute resources and is therefore largely limited to large cloud operators. Edge-assist…
Composer: A Search Framework for Hybrid Neural Architecture Design
Bilge Acun, Prasoon Sinha, Newsha Ardalani +7
Hybrid model architectures that combine computational primitives (e.g., Attention, MLP) in different ratios have shown promising performance beyond Transformers. Some studies have…
Towards Decentralized and Sustainable Foundation Model Training with the Edge
Leyang Xue, Meghana Madhyastha, Randal Burns +2
Foundation models are at the forefront of AI research, appealing for their ability to learn from vast datasets and cater to diverse tasks. Yet, their significant computational dema…
Masked Matrix Multiplication for Emergent Sparsity
Brian Wheatman, Meghana Madhyastha, Randal Burns
Artificial intelligence workloads, especially transformer models, exhibit emergent sparsity in which computations perform selective sparse access to dense data. The workloads are i…
Understanding Patterns of Deep Learning ModelEvolution in Network Architecture Search
Robert Underwood, Meghana Madhastha, Randal Burns +1
Network Architecture Search and specifically Regularized Evolution is a common way to refine the structure of a deep learning model.However, little is known about how models empiri…