1 paper
Yash Akhauri, Ahmed F AbouElhamayed, Jordan Dotzel +4
The high power consumption and latency-sensitive deployments of large language models (LLMs) have motivated efficiency techniques like quantization and sparsity. Contextual sparsit…