2 papers
cs.AR2026
Hardware Mechanisms to Dynamically Throttle AI Performance
Haiyue Ma, Lauren Malek, Joseph Forzani +1
As more capable AI models are increasingly integrated into critical computer systems, the lack of control over AI intent motivates safety mechanisms. Existing software safeguards i…
cs.AR2025
SPAD: Specialized Prefill and Decode Hardware for Disaggregated LLM Inference
Hengrui Zhang, Pratyush Patel, August Ning +1
Large Language Models (LLMs) have gained popularity in recent years, driving up the demand for inference. LLM inference is composed of two phases with distinct characteristics: a c…