50 citations · 50 across the 3 of their papers we have counts for
3 papers
HALO: Hardware-aware quantization with low critical-path-delay weights for LLM acceleration
Rohan Juneja, Shivam Aggarwal, Safeen Huda +2
Quantization is critical for efficiently deploying large language models (LLMs). Yet conventional methods remain hardware-agnostic, limited to bit-width constraints, and do not acc…
ShadowLLM: Predictor-based Contextual Sparsity for Large Language Models
Yash Akhauri, Ahmed F AbouElhamayed, Jordan Dotzel +4
The high power consumption and latency-sensitive deployments of large language models (LLMs) have motivated efficiency techniques like quantization and sparsity. Contextual sparsit…
A Full-Stack Search Technique for Domain Optimized Deep Learning Accelerators
Dan Zhang, Safeen Huda, Ebrahim Songhori +4
The rapidly-changing deep learning landscape presents a unique opportunity for building inference accelerators optimized for specific datacenter-scale workloads. We propose Full-st…