3 citations · 3 across the 5 of their papers we have counts for
Showing cs.DCShow all
2 papers · 1 filter
cs.DC2026
TaxBreak: Unmasking the Hidden Costs of LLM Inference Through Overhead Decomposition
Prabhu Vellaisamy, Shreesh Tripathi, Vignesh Natarajan +3
Large Language Model (LLM) inference is widely used in interactive assistants and agentic systems. In latency-sensitive deployments, inference time can become dominated by host-sid…
cs.DC2025
Characterizing and Optimizing LLM Inference Workloads on CPU-GPU Coupled Architectures
Prabhu Vellaisamy, Thomas Labonte, Sourav Chakraborty +3
Large language model (LLM)-based inference workloads increasingly dominate data center costs and resource utilization. Therefore, understanding the inference workload characteristi…