1 citations · 1 across the 3 of their papers we have counts for
3 papers
cs.DC2025
Utility-Driven Speculative Decoding for Mixture-of-Experts
Anish Saxena, Po-An Tsai, Hritvik Taneja +2
GPU memory bandwidth is the main bottleneck for low-latency Large Language Model (LLM) inference. Speculative decoding leverages idle GPU compute by using a lightweight drafter to…
cs.DC2024
Improving Multi-Instance GPU Efficiency via Sub-Entry Sharing TLB Design
Bingyao Li, Yueqi Wang, Tianyu Wang +4
NVIDIA's Multi-Instance GPU (MIG) technology enables partitioning GPU computing power and memory into separate hardware instances, providing complete isolation including compute re…
cs.CR2024★ 1 cited
Probabilistic Tracker Management Policies for Low-Cost and Scalable Rowhammer Mitigation
Aamer Jaleel, Stephen W. Keckler, Gururaj Saileshwar
This paper focuses on mitigating DRAM Rowhammer attacks. In recent years, solutions like TRR have been deployed in DDR4 DRAM to track aggressor rows and then issue a mitigative act…