collaborators
Showing cs.LGShow all

5 papers · 1 filter

cs.LG2025

SparAMX: Accelerating Compressed LLMs Token Generation on AMX-powered CPUs

Ahmed F. AbouElhamayed, Jordan Dotzel, Yash Akhauri +6

Large language models have high compute, latency, and memory requirements. While specialized accelerators such as GPUs and TPUs typically run these workloads, CPUs are more widely…

cs.LG2025

Low-Rank Adapters Meet Neural Architecture Search for LLM Compression

J. Pablo Muñoz, Jinjie Yuan, Nilesh Jain

The rapid expansion of Large Language Models (LLMs) has posed significant challenges regarding the computational resources required for fine-tuning and deployment. Recent advanceme…

cs.LG2025

MultiPruner: Balanced Structure Removal in Foundation Models

J. Pablo Muñoz, Jinjie Yuan, Nilesh Jain

Recently, state-of-the-art approaches for pruning large pre-trained models (LPMs) have demonstrated that the training-free removal of non-critical residual blocks in Transformers i…

cs.LG2024

SQFT: Low-cost Model Adaptation in Low-precision Sparse Foundation Models

Juan Pablo Muñoz, Jinjie Yuan, Nilesh Jain

Large pre-trained models (LPMs), such as large language models, have become ubiquitous and are employed in many applications. These models are often adapted to a desired domain or…

cs.LG2024

Shears: Unstructured Sparsity with Neural Low-rank Adapter Search

J. Pablo Muñoz, Jinjie Yuan, Nilesh Jain

Recently, several approaches successfully demonstrated that weight-sharing Neural Architecture Search (NAS) can effectively explore a search space of elastic low-rank adapters (LoR…