2 papers
cs.CL2026
Unifying Depth and Width Pruning for LLMs via Binary Knapsack Optimization
Palaash Goel, Ayan Sengupta, Akshay Nambi +1
Structured pruning is a promising approach for compressing large language models (LLMs), yet existing methods rely heavily on greedy heuristics that produce myopic decisions, and o…
cs.CL2026
It Takes a MAESTRO To Prune Bad Experts
Palaash Goel, Ayush Maheshwari, Tanmoy Chakraborty
Sparsely-activated Mixture-of-Experts (MoE) language models achieve remarkable inference efficiency by activating only a small fraction of parameters per token, yet their full expe…