1 citations · 1 across the 6 of their papers we have counts for
4 papers · 1 filter
BOOST: Concurrent Access to Host Memory and HBM to Accelerate LLM Inference
Anish Saxena, Jae Hyung Ju, Hritvik Taneja +4
GPU memory bandwidth and capacity limit throughput in large language model (LLM) inference. The GPU memory system consists of a primary tier of high-bandwidth memory (HBM) and a se…
Utility-Driven Speculative Decoding for Mixture-of-Experts
Anish Saxena, Po-An Tsai, Hritvik Taneja +2
GPU memory bandwidth is the main bottleneck for low-latency Large Language Model (LLM) inference. Speculative decoding leverages idle GPU compute by using a lightweight drafter to…
GPUArmor: A Hardware-Software Co-design for Efficient and Scalable Memory Safety on GPUs
Mohamed Tarek Ibn Ziad, Sana Damani, Mark Stephenson +2
Memory safety errors continue to pose a significant threat to current computing systems, and graphics processing units (GPUs) are no exception. A prominent class of memory safety a…
Improving Multi-Instance GPU Efficiency via Sub-Entry Sharing TLB Design
Bingyao Li, Yueqi Wang, Tianyu Wang +4
NVIDIA's Multi-Instance GPU (MIG) technology enables partitioning GPU computing power and memory into separate hardware instances, providing complete isolation including compute re…