1 citations · 1 across the 4 of their papers we have counts for
4 papers
Mamba-Shedder: Post-Transformer Compression for Efficient Selective Structured State Space Models
J. Pablo Muñoz, Jinjie Yuan, Nilesh Jain
Large pre-trained models have achieved outstanding results in sequence modeling. The Transformer block and its attention mechanism have been the main drivers of the success of thes…
Low-Rank Adapters Meet Neural Architecture Search for LLM Compression
J. Pablo Muñoz, Jinjie Yuan, Nilesh Jain
The rapid expansion of Large Language Models (LLMs) has posed significant challenges regarding the computational resources required for fine-tuning and deployment. Recent advanceme…
MultiPruner: Balanced Structure Removal in Foundation Models
J. Pablo Muñoz, Jinjie Yuan, Nilesh Jain
Recently, state-of-the-art approaches for pruning large pre-trained models (LPMs) have demonstrated that the training-free removal of non-critical residual blocks in Transformers i…
SQFT: Low-cost Model Adaptation in Low-precision Sparse Foundation Models
Juan Pablo Muñoz, Jinjie Yuan, Nilesh Jain
Large pre-trained models (LPMs), such as large language models, have become ubiquitous and are employed in many applications. These models are often adapted to a desired domain or…