3 papers
cs.DC2026
SiFAR: Synchronization-Free All-Reduce for Low-Latency LLM Inference
Hritvik Taneja, Anish Saxena, Abhishek Revinipati +3
The rise of reasoning models and agentic systems has made LLM token-generation latency a key bottleneck. Unlike chatbots, whose latency gains saturate at human reading speed, these…
cs.AR2025
LIMINAL: Exploring The Frontiers of LLM Decode Performance
Michael Davies, Neal Crago, Karthikeyan Sankaralingam +1
The rapid advancement of Large Language Models (LLMs) necessitates a deep understanding of their fundamental performance limits. This paper investigates the limits of LLM inference…
cs.AR2025
Kitsune: Enabling Dataflow Execution on GPUs
Michael Davies, Neal Crago, Karthikeyan Sankaralingam +1
State of art DL models are growing in size and complexity, with many modern models also increasing in heterogeneity of behavior. GPUs are still the dominant platform for DL applica…