4 papers
Utility-Driven Speculative Decoding for Mixture-of-Experts
Anish Saxena, Po-An Tsai, Hritvik Taneja +2
GPU memory bandwidth is the main bottleneck for low-latency Large Language Model (LLM) inference. Speculative decoding leverages idle GPU compute by using a lightweight drafter to…
GPUArmor: A Hardware-Software Co-design for Efficient and Scalable Memory Safety on GPUs
Mohamed Tarek Ibn Ziad, Sana Damani, Mark Stephenson +2
Memory safety errors continue to pose a significant threat to current computing systems, and graphics processing units (GPUs) are no exception. A prominent class of memory safety a…
QPRAC: Towards Secure and Practical PRAC-based Rowhammer Mitigation using Priority Queues
Jeonghyun Woo, Chris S. Lin, Prashant J. Nair +2
JEDEC has introduced the Per Row Activation Counting (PRAC) framework for DDR5 and future DRAMs to enable precise counting of DRAM row activations. PRAC enables a holistic mitigati…
Teaching an Old Dog New Tricks: Verifiable FHE Using Commodity Hardware
Jules Drean, Fisher Jepsen, Edward Suh +3
We present Argos, a simple approach for adding verifiability to fully homomorphic encryption (FHE) schemes using trusted hardware. Traditional approaches to verifiable FHE require…